VLDB 2026 Research / reviewers in the wild / expert
Ling-Yu Duan
dblp:d/LingyuDuan · also Lingyu Duan
· DBLP profile ↗
244ranked-venue papers
23as first author
57since 2021 · last 2026
0000-0002-4491-2023ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 210 · 22 first-author · 38 since 2021Artificial intelligence and machine learning · 64 · 31 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-author · 1 since 2021Computer networks · 7 · 4 since 2021Systems, architecture and hardware · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QuadPrior++: Multi-Dimension Augmented Physical Prior for Zero-Reference Illumination EnhancementabstractExisting low-light enhancement methods typically rely on fitting data mappings (pixel-wise mappings through fully supervised methods or distribution-wise mappings through weakly supervised or self-supervised methods). However, their performance is heavily dependent on specific scenes and fails to adequately model the intrinsic prior of natural images, resulting in poor generalization. To tackle this challenge, we leverage the strengths of powerful generative diffusion models, conditioned on a thoughtfully designed prior, and propose a novel zero-reference low-light enhancement framework that gets rid of dependence on the distribution of low-light images. In detail, we address the most fundamental core by proposing an illumination-invariant prior derived from the theory of physical light transfer, bridging the gap between normal and low-light domains, and enabling zero-shot enhancement without the need for low-light-specific training. A prior-to-image restoration framework is built upon generative diffusion models, pre-trained on normal-light data. During inference, the framework extracts the illumination-invariant prior from low-light inputs and maps them back to high-quality images, naturally for low-light enhancement. Additionally, such intrinsic properties of illumination-invariant prior open up opportunities for distilling diffusion models into compact CNN-based networks. We propose a novel prior-injected distillation paradigm incorporating intensity, frequency, and gradient domain-augmented regularization comprehensively. This distillation framework not only reduces computational costs but also maintains high fidelity and perceptual quality in enhanced outputs, making it more efficient and practical for real-world applications. The approach further extends seamlessly to handle over-exposure scenarios, demonstrating its versatility in addressing complex lighting conditions. Extensive experiments demonstrate the superiority of our framework in various scenarios, as well as its strong interpretability, robustness, and efficiency. Haofeng Huang, Wenjing Wang 0001, Wenhan Yang, Ling-Yu Duan, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Local Dimension Enhancement Representation Learning for Skeleton-Based Action SegmentationabstractMost existing self-supervised learning methods for skeleton-based temporal action segmentation (TAS) fail to capture the short-term motion semantics essential for dense frame-level prediction, as they typically learn representations that are either too coarse or motion-insensitive. This issue is reflected in local dimension collapse, which highlights the limitations of current approaches and suggests directions for improvement. Specifically, to address the issue of local dimension collapse for self-supervised learning in TAS, we propose the Local Dimension Enhancement (LoDE) framework, which introduces the local effective rank (LER) as a metric to measure and a learning objective to reduce this collapse. A new fine-grained representation scale, termed a motion unit, is defined as a temporal clip of consecutive skeleton frames to model skeleton data. Centered on this representation scale, we analyze existing methods (sequence-scale and frame-scale learning) with the tool of LER and theoretically demonstrate that introducing motion unit-scale learning is essential to alleviate local dimension collapse. Inspired by our theoretical insights, we design a multi-scale semantics module that integrates frame-, sequence-, and motion unit-scale learning, with LER-based regularization to enrich local representation diversity. These designs effectively alleviate local dimension collapse and lead to significant improvements in TAS, as evidenced by LoDE's superior performance over state-of-the-art methods on three large-scale untrimmed datasets: PKUMMD, TSU, and BABEL. Our project website is available at https://carefreesun.github.io/LoDE_TIP_2026/. Shaofan Sun, Lilang Lin, Jiahang Zhang 0001, Ling-Yu Duan, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | Enabling Real-World Supervised Video Anomaly Detection: New Open-Set Benchmark and New FrameworkabstractThe inherent unpredictability of abnormal events in real-world Video Anomaly Detection (VAD) presents significant challenges for model generalization. Early unsupervised methods aim to detect any anomalies as deviations from normal patterns, but they often equate rarity with abnormality, an assumption that leads to high false positives in open-world scenarios where uncommon behaviors are not necessarily abnormal. Current mainstream supervised approaches focus on memorizing discriminative features rather than learning the underlying pattern that governs normal and abnormal events. Although their performance on the seen anomalies is superior, their capability for robust open-set inference is still limited. Moreover, conventional closed-set evaluation benchmarks obscure this critical distinction by assuming identical anomaly types during training and testing. To address this limitation, this paper constructs the Real-world Supervised Open-set Benchmark (RSOB) built on a novel large-scale traffic abnormal event dataset with precise frame-level annotation, supporting evaluation across varying supervision granularities, from full to weak supervision settings. This benchmark is also the first to enable the real-world open-set evaluation constructed on the simulation of potential open-set scenarios in the application. Moreover, we further propose a human-prior-aware framework that learns domain-agnostic normality rules through a novel Human-Prior-Focused (HPF) feature space. This space is derived from a semantic-aware transformation on pre-trained feature space to effectively separate normal events while modeling generalizable human rules. Notably, the framework is architecture-agnostic and establishes the first unified solution applicable to both supervised and weakly supervised VAD paradigms. Our extensive experiments demonstrate the superiority of our framework over traditional discriminative methods across our benchmark as well as the conventional benchmark Ubnormal. The code and dataset will be publicly available. Zhuo Chen 0006, Ling-Yu Duan |
IEEE Trans. Multim. | 7 |
| 2026 | Seeing in the Dark with Ambient GuidanceabstractA low-light image taken in a dark scene usually suffers from severe distortions, which does not accurately characterize the ambient lighting. Long exposure is an accustomed way to capture more supplementary light and alleviate the degradation, but sometimes it induces other distortions, e.g., blurriness. To address this issue, we propose a new paradigm that introduces additional captured ambient guidance, i.e., a long-exposure image to steer the low-light enhancement. In practice, this long-exposure image can be obtained conveniently, but usually suffers from blurriness and misalignment. To effectively extract and fuse information from degraded and misaligned low-light and guidance image pairs, we propose a Long Exposure Compensation Network (LECNet). Adaptive Band Regression is introduced to disentangle the image into multi-scale representations and coarse-to-fine aggregate them with an attention mechanism. For stable image-guidance registration and artifact suppression, we propose a Bounded Cross-domain Deformable Alignment to warp the guidance based on extracted feature pyramids step by step. To integrate knowledge about the degradation into our LECNet for better fidelity, a dual learned back projection is enforced between the predicted result and the paired inputs in illumination and texture detail consistency, serving the model training for both offline training and online sample-adaptive finetuning. For training and evaluation of this new paradigm, we build a dataset with both synthetic and real-captured image triplets of long/short exposure pairs and extra blurry guidance. The experimental evaluation demonstrates the significance of our new paradigm, as well as the superiority of our LECNet and its usability in the real world. Haofeng Huang, Wenhan Yang, Mengnan Wang, Ling-Yu Duan, Jiaying Liu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Theoretical Insights in Model Inversion Robustness and Conditional Entropy Maximization for Collaborative Inference SystemsabstractBy locally encoding raw data into intermediate features, collaborative inference enables end users to leverage powerful deep learning models without exposure of sensitive raw data to cloud servers. However, recent studies have revealed that these intermediate features may not sufficiently preserve privacy, as information can be leaked and raw data can be reconstructed via model inversion attacks (MIAs). Obfuscation-based methods, such as noise corruption, adversarial representation learning, and information filters, enhance the inversion robustness by obfuscating the task-irrelevant redundancy empirically. However, methods for quantifying such redundancy remain elusive, and the explicit mathematical relation between this redundancy minimization and inversion robustness enhancement has not yet been established. To address that, this work first theoretically proves that the conditional entropy of inputs given intermediate features provides a guaranteed lower bound on the reconstruction mean square error (MSE) under any MIA. Then, we derive a differentiable and solvable measure for bounding this conditional entropy based on the Gaussian mixture estimation and propose a conditional entropy maximization (CEM) algorithm to enhance the inversion robustness. Experimental results on four datasets demonstrate the effectiveness and adaptability of our proposed CEM; without compromising feature utility and computing efficiency, plugging the proposed CEM into obfuscation-based defense mechanisms consistently boosts their inversion robustness, achieving average gains ranging from 12.9% to 48.2%. Code is available at https://github.com/xiasong0501/CEM. Song Xia, Yi Yu 0011, Wenhan Yang, Meiwen Ding, Zhuo Chen 0006, Ling-Yu Duan, Alex Chichung Kot, Xudong Jiang 0001 |
CVPR | 6 |
| 2025 | Adaptive Dual Uncertainty Optimization: Boosting Monocular 3D Object Detection under Test-Time ShiftsabstractAccurate monocular 3D object detection (M3OD) is pivotal for safety-critical applications like autonomous driving, yet its reliability deteriorates significantly under real-world domain shifts caused by environmental or sensor variations. To address these shifts, Test-Time Adaptation (TTA) methods have emerged, enabling models to adapt to target distributions during inference. While prior TTA approaches recognize the positive correlation between low uncertainty and high generalization ability, they fail to address the dual uncertainty inherent to M3OD: semantic uncertainty (ambiguous class predictions) and geometric uncertainty (unstable spatial localization). To bridge this gap, we propose Dual Uncertainty Optimization (DUO), the first TTA framework designed to jointly minimize both uncertainties for robust M3OD. Through a convex optimization lens, we introduce an innovative convex structure of the focal loss and further derive a novel unsupervised version, enabling label-agnostic uncertainty weighting and balanced learning for high-uncertainty objects. In parallel, we design a semantic-aware normal field constraint that preserves geometric coherence in regions with clear semantic cues, reducing uncertainty from the unstable 3D representation. This dual-branch mechanism forms a complementary loop: enhanced spatial perception improves semantic classification, and robust semantic predictions further refine spatial understanding. Extensive experiments demonstrate the superiority of DUO over existing methods across various datasets and domain shift types. Xinzhu Ma, Shixiang Tang, Wenhan Yang, Ling-Yu Duan |
ICCV | 7 |
| 2025 | Which Tasks Should Be Compressed Together? A Causal Discovery Approach for Efficient Multi-Task Representation CompressionabstractConventional image compression methods are inadequate for intelligent analysis, as they overemphasize pixel-level precision while neglecting semantic significance and the interaction among multiple tasks. This paper introduces a Taskonomy-Aware Multi-Task Compression framework comprising (1) inter-coherent task grouping, which organizes synergistic tasks into shared representations to improve multi-task accuracy and reduce encoding volume, and (2) a conditional entropy-based directed acyclic graph (DAG) that captures causal dependencies among grouped representations. By leveraging parent representations as contextual priors for child representations, the framework effectively utilizes cross-task information to improve entropy model accuracy. Experiments on diverse vision tasks, including Keypoint 2D, Depth Z-buffer, Semantic Segmentation, Surface Normal, Edge Texture, and Autoencoder, demonstrate significant bitrate-performance gains, validating the method’s capability to reduce system entropy uncertainty. These findings underscore the potential of leveraging representation disentanglement, synergy, and causal modeling to learn compact representations, which enable efficient multi-task compression in intelligent systems. Sha Guo, Zhuo Chen 0006, Wenhan Yang, Ling-Yu Duan |
ICLR | 8 |
| 2025 | Adaptive Gradient Quantization with Bit Allocation for Distributed Deep LearningabstractGradient compression plays a crucial role in mitigating communication overhead in distributed deep learning. Existing gradient compression methods usually employ fix-bit quantization across all layers, neglecting the varying sensitivities of different layers to compression, resulting in suboptimal performance. In this paper, we introduce a layer-wise bit allocation mechanism for gradient quantization that minimizes overall quantization error within a specified bit budget. To address the heavy computational load of conventional greedy search approach for bit allocation, we develop two acceleration techniques to reduce computational overhead, thereby making the proposed bit allocation method feasible for real-time deep learning training. Specifically, by observing the bit allocation statistics, we propose Bit Searching Range Optimization to narrow the available bit options, while the Bit Pre-Assignment selectively bypasses certain searching processes. Experimental results across various neural network models and datasets demonstrate the effectiveness of our proposed bit allocation methods for gradient quantization. The combination of proposed acceleration techniques offers an advantageous trade-off among quantization error, model training performance and time consumption. Moreover, our proposed bit allocation methods can be seamlessly integrated with existing gradient compression approaches, improving overall performance. Fei Gao 0019, Wenhan Yang, Ling-Yu Duan, Zhuo Chen 0006 |
ICME | 5 |
| 2025 | Beyond Entropy: Region Confidence Proxy for Wild Test-Time AdaptationabstractWild Test-Time Adaptation (WTTA) is proposed to adapt a source model to unseen domains under extreme data scarcity and multiple shifts. Previous approaches mainly focused on sample selection strategies, while overlooking the fundamental problem on underlying optimization. Initially, we critically analyze the widely-adopted entropy minimization framework in WTTA and uncover its significant limitations in noisy optimization dynamics that substantially hinder adaptation efficiency. Through our analysis, we identify region confidence as a superior alternative to traditional entropy, however, its direct optimization remains computationally prohibitive for real-time applications. In this paper, we introduce a novel region-integrated method **ReCAP** that bypasses the lengthy process. Specifically, we propose a probabilistic region modeling scheme that flexibly captures semantic changes in embedding space. Subsequently, we develop a finite-to-infinite asymptotic approximation that transforms the intractable region confidence into a tractable and upper-bounded proxy. These innovations significantly unlock the overlooked potential dynamics in local region in a concise solution. Our extensive experiments demonstrate the consistent superiority of ReCAP over existing methods across various datasets and wild scenarios. The source code will be available at https://github.com/hzcar/ReCAP. Yichun Hu, Shixiang Tang, Ling-Yu Duan |
ICML | 5 |
| 2025 | LabUtopia: High-Fidelity Simulation and Hierarchical Benchmark for Scientific Embodied AgentsabstractScientific embodied agents play a crucial role in modern laboratories by automating complex experimental workflows.Compared to typical household environments, laboratory settings impose significantly higher demands on perception of physical-chemical transformations and long-horizon planning, making them an ideal testbed for advancing embodied intelligence.However, its development has been long hampered by the lack of suitable simulator and benchmarks.In this paper, we address this gap by introducing LabUtopia, a comprehensive simulation and benchmarking suite designed to facilitate the development of generalizable, reasoning-capable embodied agents in laboratory settings. Specifically, it integrates i) LabSim, a high-fidelity simulator supporting multi-physics and chemically meaningful interactions; ii) LabScene, a scalable procedural generator for diverse scientific scenes; and iii) LabBench, a hierarchical benchmark spanning five levels of complexity from atomic actions to long-horizon mobile manipulation. LabUtopia supports 30 distinct tasks and includes more than 200 scene and instrument assets, enabling large-scale training and principled evaluation in high-complexity environments.We demonstrate that LabUtopia offers a powerful platform for advancing the integration of perception, planning, and control in scientific-purpose agents and provides a rigorous testbed for exploring the practical capabilities and generalization limits of embodied intelligence in future research. Project web page: https://rui-li023.github.io/labutopia-site/ Rui Li 0054, Wenxi Qu, Jinouwen Zhang, Zhenfei Yin, Sha Zhang 0002, Xuantuo Huang, Jiangmiao Pang, Wanli Ouyang, Lei Bai 0001, Wangmeng Zuo, Ling-Yu Duan, Dongzhan Zhou, Shixiang Tang |
NeurIPS | 14 |
| 2025 | Bridging the Source-to-Target Gap for Cross-Domain Person Re-identification with Intermediate Domains
Yongxing Dai, Yifan Sun 0003, Jun Liu 0036, Zekun Tong, Ling-Yu Duan |
Int. J. Comput. Vis. | 5 |
| 2025 | DM-PCL: Text-Driven Dual-Modal Prototype Consistency Learning for Weakly-Supervised Few-Shot Part Segmentation
Mengya Han, Yong Luo 0002, Han Hu 0003, Zengmao Wang, Lefei Zhang, Bo Du 0001, Ling-Yu Duan, Dacheng Tao |
Int. J. Comput. Vis. | 7 |
| 2025 | Robust and Transferable Backdoor Attacks Against Deep Image Compression With Selective Frequency PriorabstractRecent advancements in deep learning-based compression techniques have demonstrated remarkable performance surpassing traditional methods. Nevertheless, deep neural networks have been observed to be vulnerable to backdoor attacks, where an added pre-defined trigger pattern can induce the malicious behavior of the models. In this paper, we propose a novel approach to launch a backdoor attack with multiple triggers against learned image compression models. Drawing inspiration from the widely used discrete cosine transform (DCT) in existing compression codecs and standards, we propose a frequency-based trigger injection model that adds triggers in the DCT domain. In particular, we design several attack objectives that are adapted for a series of diverse scenarios, including: 1) attacking compression quality in terms of bit-rate and reconstruction quality; 2) attacking task-driven measures, such as face recognition and semantic segmentation in downstream applications. To facilitate more efficient training, we develop a dynamic loss function that dynamically balances the impact of different loss terms with fewer hyper-parameters, which also results in more effective optimization of the attack objectives with improved performance. Furthermore, we consider several advanced scenarios. We evaluate the resistance of the proposed backdoor attack to the defensive pre-processing methods and then propose a two-stage training schedule along with the design of robust frequency selection to further improve resistance. To strengthen both the cross-model and cross-domain transferability on attacking downstream CV tasks, we propose to shift the classification boundary in the attack loss during training. Extensive experiments also demonstrate that by employing our trained trigger injection models and making slight modifications to the encoder parameters of the compression model, our proposed attack can successfully inject multiple backdoors accompanied by their corresponding triggers into a single image compression model. Yi Yu 0011, Yufei Wang 0006, Wenhan Yang, Lanqing Guo, Shijian Lu, Ling-Yu Duan, Yap-Peng Tan, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Seeing Dark Videos via Self-Learned Bottleneck Neural RepresentationabstractEnhancing low-light videos in a supervised style presents a set of challenges, including limited data diversity, misalignment, and the domain gap introduced through the dataset construction pipeline. Our paper tackles these challenges by constructing a self-learned enhancement approach that gets rid of the reliance on any external training data. The challenge of self-supervised learning lies in fitting high-quality signal representations solely from input signals. Our work designs a bottleneck neural representation mechanism that extracts those signals. More in detail, we encode the frame-wise representation with a compact deep embedding and utilize a neural network to parameterize the video-level manifold consistently. Then, an entropy constraint is applied to the enhanced results based on the adjacent spatial-temporal context to filter out the degraded visual signals, e.g. noise and frame inconsistency. Last, a novel Chromatic Retinex decomposition is proposed to effectively align the reflectance distribution temporally. It benefits the entropy control on different components of each frame and facilitates noise-to-noise training, successfully suppressing the temporal flicker. Extensive experiments demonstrate the robustness and superior effectiveness of our proposed method. Our project is publicly available at: https://huangerbai.github.io/SLBNR/. Haofeng Huang, Wenhan Yang, Ling-Yu Duan, Jiaying Liu 0001 |
AAAI | 3 |
| 2024 | Evidential Uncertainty-Guided Mitochondria Segmentation for 3D EM ImagesabstractRecent advances in deep learning have greatly improved the segmentation of mitochondria from Electron Microscopy (EM) images. However, suffering from variations in mitochondrial morphology, imaging conditions, and image noise, existing methods still exhibit high uncertainty in their predictions. Moreover, in view of our findings, predictions with high levels of uncertainty are often accompanied by inaccuracies such as ambiguous boundaries and amount of false positive segments. To deal with the above problems, we propose a novel approach for mitochondria segmentation in 3D EM images that leverages evidential uncertainty estimation, which for the first time integrates evidential uncertainty to enhance the performance of segmentation. To be more specific, our proposed method not only provides accurate segmentation results, but also estimates associated uncertainty. Then, the estimated uncertainty is used to help improve the segmentation performance by an uncertainty rectification module, which leverages uncertainty maps and multi-scale information to refine the segmentation. Extensive experiments conducted on four challenging benchmarks demonstrate the superiority of our proposed method over existing approaches. Ruohua Shi, Ling-Yu Duan, Tiejun Huang 0001, Tingting Jiang 0001 |
AAAI | 2 |
| 2024 | LEAD: Exploring Logit Space Evolution for Model SelectionabstractThe remarkable success of “pretrain-then-finetune” paradigm has led to a proliferation of available pre-trained models for vision tasks. This surge presents a significant challenge in efficiently choosing the most suitable pre-trained models for downstream tasks. The critical aspect of this challenge lies in effectively predicting the model transferability by considering the underlying fine-tuning dynamics. Existing methods often model fine-tuning dynamics in feature space with linear transformations, which do not precisely align with the fine-tuning objective and fail to grasp the essential nonlinearity from optimization. To this end, we present LEAD, a finetuning-aligned approach based on the network output of logits. LEAD proposes a theoretical framework to model the optimization process and derives an ordinary differential equation (ODE) to depict the nonlinear evolution toward the final logit state. Additionally, we design a class-aware decomposition method to consider the varying evolution dynamics across classes and further ensure practical applicability. Integrating the closely aligned optimization objective and nonlinear modeling capabilities derived from the differential equation, our method offers a concise solution to effectively bridge the optimization gap in a single step, bypassing the lengthy fine-tuning process. The comprehensive experiments on 24 supervised and self-supervised pre-trained models across 10 downstream datasets demonstrate impressive performances and showcase its broad adaptability even in low-data scenarios. Shixiang Tang, Jun Liu 0036, Yichun Hu, Ling-Yu Duan |
CVPR | 6 |
| 2024 | A Unified Image Compression Method for Human Perception and Multiple Vision Tasks
Sha Guo, Lin Sui, Chen-Lin Zhang, Zhuo Chen 0006, Wenhan Yang, Ling-Yu Duan |
ECCV (71) | 6 |
| 2024 | ShapeMamba-EM: Fine-Tuning Foundation Model with Local Shape Descriptors and Mamba Blocks for 3D EM Image Segmentation
Ruohua Shi, Qiufan Pang, Lei Ma 0008, Ling-Yu Duan, Tiejun Huang 0001, Tingting Jiang 0001 |
MICCAI (12) | 4 |
| 2024 | DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal PerceptionabstractExisting Multimodal Large Language Models (MLLMs) increasingly emphasize complex understanding of various visual elements, including multiple objects, text information, spatial relations. Their development for comprehensive visual perception hinges on the availability of high-quality image-text datasets that offer diverse visual elements and throughout image descriptions. However, the scarcity of such hyper-detailed datasets currently hinders progress within the MLLM community. The bottleneck stems from the limited perceptual capabilities of current caption engines, which fall short in providing complete and accurate annotations. To facilitate the cutting-edge research of MLLMs on comprehensive vision perception, we thereby propose Perceptual Fusion, using a low-budget but highly effective caption engine for complete and accurate image descriptions. Specifically, Perceptual Fusion integrates diverse perception experts as image priors to provide explicit information on visual elements and adopts an efficient MLLM as a centric pivot to mimic advanced MLLMs' perception abilities. We carefully select 1M highly representative images from uncurated LAION dataset and generate dense descriptions using our engine, dubbed DenseFusion-1M. Extensive experiments validate that our engine outperforms its counterparts, where the resulting dataset significantly improves the perception and cognition abilities of existing MLLMs across diverse vision-language benchmarks, especially with high-resolution images as inputs. The code and dataset are available at https://huggingface.co/datasets/BAAI/DenseFusion-1M. Haiwen Diao, Yueze Wang, Ling-Yu Duan |
NeurIPS | 6 |
| 2024 | Transferable Adversarial Attacks on SAM and Its Downstream ModelsabstractThe utilization of large foundational models has a dilemma: while fine-tuning downstream tasks from them holds promise for making use of the well-generalized knowledge in practical applications, their open accessibility also poses threats of adverse usage.
This paper, for the first time, explores the feasibility of adversarial attacking various downstream models fine-tuned from the segment anything model (SAM), by solely utilizing the information from the open-sourced SAM.
In contrast to prevailing transfer-based adversarial attacks, we demonstrate the existence of adversarial dangers even without accessing the downstream task and dataset to train a similar surrogate model.
To enhance the effectiveness of the adversarial attack towards models fine-tuned on unknown datasets, we propose a universal meta-initialization (UMI) algorithm to extract the intrinsic vulnerability inherent in the foundation model, which is then utilized as the prior knowledge to guide the generation of adversarial perturbations.
Moreover, by formulating the gradient difference in the attacking process between the open-sourced SAM and its fine-tuned downstream models, we theoretically demonstrate that a deviation occurs in the adversarial update direction by directly maximizing the distance of encoded feature embeddings in the open-sourced SAM.
Consequently, we propose a gradient robust loss that simulates the associated uncertainty with gradient-based noise augmentation to enhance the robustness of generated adversarial examples (AEs) towards this deviation, thus improving the transferability.
Extensive experiments demonstrate the effectiveness of the proposed universal meta-initialized and gradient robust adversarial attack (UMI-GRAT) toward SAMs and their downstream models.
Code is available at https://github.com/xiasong0501/GRAT. Song Xia, Wenhan Yang, Yi Yu 0011, Xun Lin, Henghui Ding, Ling-Yu Duan, Xudong Jiang 0001 |
NeurIPS | 6 |
| 2024 | Video Coding for Machines: Compact Visual Representation Compression for Intelligent Collaborative AnalyticsabstractAs an emerging research practice leveraging recent advanced AI techniques, e.g. deep models based prediction and generation, Video Coding for Machines (VCM) is committed to bridging to an extent separate research tracks of video/image compression and feature compression, and attempts to optimize compactness and efficiency jointly from a unified perspective of high accuracy machine vision and full fidelity human vision. With the rapid advances of deep feature representation and visual data compression in mind, in this paper, we summarize VCM methodology and philosophy based on existing academia and industrial efforts. The development of VCM follows a general rate-distortion optimization, and the categorization of key modules or techniques is established including feature-assisted coding, scalable coding, intermediate feature compression/optimization, and machine vision targeted codec, from broader perspectives of vision tasks, analytics resources, etc. From previous works, it is demonstrated that, although existing works attempt to reveal the nature of scalable representation in bits when dealing with machine and human vision tasks, there remains a rare study in the generality of low bit rate representation, and accordingly how to support a variety of visual analytic tasks. Therefore, we investigate a novel visual information compression for the analytics taxonomy problem to strengthen the capability of compact visual representations extracted from multiple tasks for visual analytics. A new perspective of task relationships versus compression is revisited. By keeping in mind the transferability among different machine vision tasks (e.g. high-level semantic and mid-level geometry-related), we aim to support multiple tasks jointly at low bit rates. In particular, to narrow the dimensionality gap between neural network generated features extracted from pixels and a variety of machine vision features/labels (e.g. scene class, segmentation labels), a codebook hyperprior is designed to compress the neural network-generated features. As demonstrated in our experiments, this new hyperprior model is expected to improve feature compression efficiency by estimating the signal entropy more accurately, which enables further investigation of the granularity of abstracting compact features among different tasks. Wenhan Yang, Haofeng Huang, Yueyu Hu, Ling-Yu Duan, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | HARDer-Net: Hardness-Guided Discrimination Network for 3D Early Activity PredictionabstractTo predict the class label from a partially observable activity sequence can be quite challenging due to the high degree of similarity existing in early segments of different activities. In this paper, an innovative HARDness-Guided Discrimination Network (HARDer-Net) is proposed to evaluate the relationship between similar activity pairs that are extremely hard to discriminate. To train our HARDer-Net, an innovative adversarial learning scheme has been designed, providing our network with the strength to extract subtle discrimination information for the prediction of 3D early activities. Moreover, to enhance the adversarial learning scheme efficacy of our model for 3D early action prediction, we construct a Hardness-Guided bank that dynamically records the hard similar samples and conducts reward-guided selections of these recorded hard samples using a deep reinforcement learning scheme. The proposed method significantly enhances the capability of the model to discern fine-grained differences in early activity sequences. Several widely-used activity datasets are used to evaluate our proposed HARDer-Net, and we achieve state-of-the-art performance across all the evaluated datasets. Wei Zhang 0021, Ling-Yu Duan, Jun Liu 0036 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Switchable Representation Learning Framework with Self-CompatibilityabstractReal-world visual search systems involve deployments on multiple platforms with different computing and storage resources. Deploying a unified model that suits the minimal-constrain platforms leads to limited accuracy. It is expected to deploy models with different capacities adapting to the resource constraints, which requires features extracted by these models to be aligned in the metric space. The method to achieve feature alignments is called “compatible learning”. Existing research mainly focuses on the one-to-one compatible paradigm, which is limited in learning compatibility among multiple models. We propose a Switchable representation learning Framework with Self-Compatibility (SFSC). SFSC generates a series of compatible sub-models with different capacities through one training process. The optimization of sub-models faces gradients conflict, and we mitigate this problem from the perspective of the magnitude and direction. We adjust the priorities of sub-models dynamically through uncertainty estimation to co-optimize sub-models properly. Besides, the gradients with conflicting directions are projected to avoid mutual interference. SFSC achieves state-of-the-art performance on the evaluated datasets. Shengsen Wu, Yihang Lou, Xiongkun Linghu, Ling-Yu Duan |
CVPR | 6 |
| 2023 | Exploring Model Transferability through the Lens of Potential EnergyabstractTransfer learning has become crucial in computer vision tasks due to the vast availability of pre-trained deep learning models. However, selecting the optimal pre-trained model from a diverse pool for a specific downstream task remains a challenge. Existing methods for measuring the transferability of pre-trained models rely on statistical correlations between encoded static features and task labels, but they overlook the impact of underlying representation dynamics during fine-tuning, leading to unreliable results, especially for self-supervised models. In this paper, we present an insightful physics-inspired approach named PED to address these challenges. We reframe the challenge of model selection through the lens of potential energy and directly model the interaction forces that influence fine-tuning dynamics. By capturing the motion of dynamic representations to decline the potential energy within a force-driven physical model, we can acquire an enhanced and more stable observation for estimating transferability. The experimental results on 10 downstream tasks and 12 self-supervised models demonstrate that our approach can seamlessly integrate into existing ranking techniques and enhance their performances, revealing its effectiveness for the model selection task and its potential for understanding the mechanism in transfer learning. Code is available at https://github.com/lixiaotong97/PED. Yixiao Ge, Ying Shan, Ling-Yu Duan |
ICCV | 5 |
| 2023 | Toward Scalable Image Feature Compression: A Content-Adaptive and Diffusion-Based ApproachabstractTraditional image codecs prioritize signal fidelity and human perception, often neglecting machine vision tasks. Deep learning approaches have shown promising coding performance by leveraging rich semantic embeddings that can be optimized for both human and machine vision. However, these compact embeddings struggle to represent low-level details like contours and textures, leading to imperfect reconstructions. Additionally, existing learning-based coding tools lack scalability. To address these challenges, this paper presents a content-adaptive diffusion model for scalable image compression. The method encodes accurate texture through a diffusion process, enhancing human perception while preserving important features for machine vision tasks. It employs a Markov palette diffusion model with commonly-used feature extractors and image generators, enabling efficient data compression. By utilizing collaborative texture-semantic feature extraction and pseudo-label generation, the approach accurately learns texture information. A content-adaptive Markov palette diffusion model is then applied to capture both low-level texture and high-level semantic knowledge in a scalable manner. This framework enables elegant compression ratio control by flexibly selecting intermediate diffusion states, eliminating the need for deep learning model re-training at different operating points. Extensive experiments demonstrate the effectiveness of the proposed framework in image reconstruction and downstream machine vision tasks such as object detection, segmentation, and facial landmark detection. It achieves superior perceptual quality scores compared to state-of-the-art methods. Sha Guo, Zhuo Chen 0006, Yang Zhao 0002, Ning Zhang 0023, Ling-Yu Duan |
ACM Multimedia | 6 |
| 2023 | PS-Net: human perception-guided segmentation network for EM cell membraneabstractMOTIVATION: Cell membrane segmentation in electron microscopy (EM) images is a crucial step in EM image processing. However, while popular approaches have achieved performance comparable to that of humans on low-resolution EM datasets, they have shown limited success when applied to high-resolution EM datasets. The human visual system, on the other hand, displays consistently excellent performance on both low and high resolutions. To better understand this limitation, we conducted eye movement and perceptual consistency experiments. Our data showed that human observers are more sensitive to the structure of the membrane while tolerating misalignment, contrary to commonly used evaluation criteria. Additionally, our results indicated that the human visual system processes images in both global-local and coarse-to-fine manners. RESULTS: Based on these observations, we propose a computational framework for membrane segmentation that incorporates these characteristics of human perception. This framework includes a novel evaluation metric, the perceptual Hausdorff distance (PHD), and an end-to-end network called the PHD-guided segmentation network (PS-Net) that is trained using adaptively tuned PHD loss functions and a multiscale architecture. Our subjective experiments showed that the PHD metric is more consistent with human perception than other criteria, and our proposed PS-Net outperformed state-of-the-art methods on both low- and high-resolution EM image datasets as well as other natural image datasets. AVAILABILITY AND IMPLEMENTATION: The code and dataset can be found at https://github.com/EmmaSRH/PS-Net. Ruohua Shi, Keyan Bi, Lei Ma 0008, Fang Fang 0003, Ling-Yu Duan, Tingting Jiang 0001, Tiejun Huang 0001 |
Bioinform. | 6 |
| 2023 | Benchmarking Single-Image Reflection Removal AlgorithmsabstractReflection removal has been discussed for more than decades. This paper aims to provide the analysis for different reflection properties and factors that influence image formation, an up-to-date taxonomy for existing methods, a benchmark dataset, and the unified benchmarking evaluations for state-of-the-art (especially learning-based) methods. Specifically, this paper presents a SIngle-image Reflection Removal Plus dataset “SIR$^{2+}$” with the new consideration for in-the-wild scenarios and glass with diverse color and unplanar shapes. We further perform quantitative and visual quality comparisons for state-of-the-art single-image reflection removal algorithms. Open problems for improving reflection removal algorithms are discussed at the end. Our dataset and follow-up update can be found athttps://reflectionremoval.github.io/sir2data/. Renjie Wan, Boxin Shi, Haoliang Li, Yuchen Hong, Ling-Yu Duan, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Coarse-to-fine Disentangling Demoiréing Framework for Recaptured Screen ImagesabstractRemoving the undesired moiré patterns from images capturing the contents displayed on screens is of increasing research interest, as the need for recording and sharing the instant information conveyed by the screens is growing. Previous demoiréing methods provide limited investigations into the formation process of moiré patterns to exploit moiré-specific priors for guiding the learning of demoiréing models. In this paper, we investigate the moiré pattern formation process from the perspective of signal aliasing, and correspondingly propose a coarse-to-fine disentangling demoiréing framework. In this framework, we first disentangle the moiré pattern layer and the clean image with alleviated ill-posedness based on the derivation of our moiré image formation model. Then we refine the demoiréing results exploiting both the frequency domain features and edge attention, considering moiré patterns' property on spectrum distribution and edge intensity revealed in our aliasing based analysis. Experiments on several datasets show that the proposed method performs favorably against state-of-the-art methods. Besides, the proposed method is validated to adapt well to different data sources and scales, especially on the high-resolution moiré images. Ce Wang 0007, Shengsen Wu, Renjie Wan, Boxin Shi, Ling-Yu Duan |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Dual-Tuning: Joint Prototype Transfer and Structure Regularization for Compatible Feature LearningabstractVisual retrieval system faces frequent model update and deployment. It is a heavy workload to re-extract features of the whole database every time. Feature compatibility enables the learned new visual features to be directly compared with the old features stored in the database. In this way, when updating the deployed model, we can bypass the inflexible and time-consuming feature re-extraction process. However, the old feature space that needs to be compatible is not ideal and faces outlier samples. Besides, the new and old models may be supervised by different losses, which will further causes distribution discrepancy problem between these two feature spaces. In this article, we propose a global optimization Dual-Tuning method to obtain feature compatibility against different networks and losses. A feature-level prototype loss is proposed to explicitly align two types of embedding features, by transferring global prototype information. Furthermore, we design a component-level mutual structural regularization to implicitly optimize the feature intrinsic structure. Experiments are conducted on six datasets, including person ReID datasets, face recognition datasets, and million-scale ImageNet and Place365. Experimental results demonstrate that our Dual-Tuning is able to obtain feature compatibility without sacrificing performance. Jile Jiao, Yihang Lou, Shengsen Wu, Jun Liu 0036, Xuetao Feng, Ling-Yu Duan |
IEEE Trans. Multim. | 7 |
| 2023 | Purifying Low-Light Images via Near-Infrared Enlightened ImageabstractCameras usually produce low-quality images under low-light conditions. Though many methods have been proposed to enhance the visibility of low-light images, they are mainly designed for illumination correction and less capable of sup-pressing the artifacts. In this paper, we propose to enhance the visibility and suppress artifacts by purifying low-light images under the guidance of the NIR enlightened image captured by using the near-infrared light as compensation. Specifically, we introduce a disentanglement framework to disentangle the structure and color components from the NIR enlightened and RGB images, respectively. Correspondingly, we introduce a new dataset with the RGB and NIR enlightened images for training and evaluation purposes. The experimental results show that our proposed method achieves promising results. Renjie Wan, Boxin Shi, Wenhan Yang, Bihan Wen, Ling-Yu Duan, Alex Chichung Kot |
IEEE Trans. Multim. | 5 |
| 2023 | Background Scene Recovery From an Image Looking Through Colored GlassabstractColored glass, which is commonly seen in modern city life, often degrades images taken through it with co-occurring reflection and color bias due to its optical property of simultaneous transmission, reflection, and wavelength-selective absorption. Recovering the clean background behind colored glass is inherently challenging due to the mutual interference of two degradations within a single mixture observation, and has barely been specifically considered by existing image restoration methods. In this paper, we aim at realizing faithful background scene recovery for an image taken in front of colored glass. We first analyze the formation model of mixed degradations caused by colored glass, and propose a cooperative framework to address the mutual interference problem, featuring a novel glass color invariant loss and progressive refinement. Besides, we propose a data synthesis strategy for network training. Experimental results on our newly collected real-world dataset show that our proposed method achieves state-of-the-art performance. Ce Wang 0007, Dejia Xu, Renjie Wan, Boxin Shi, Ling-Yu Duan |
IEEE Trans. Multim. | 6 |
| 2022 | Neighborhood Consensus Contrastive Learning for Backward-Compatible RepresentationabstractIn object re-identification (ReID), the development of deep learning techniques often involves model updates and deployment. It is unbearable to re-embedding and re-index with the system suspended when deploying new models. Therefore, backward-compatible representation is proposed to enable ``new'' features to be compared with ``old'' features directly, which means that the database is active when there are both ``new'' and ``old'' features in it. Thus we can scroll-refresh the database or even do nothing on the database to update. The existing backward-compatible methods either require a strong overlap between old and new training data or simply conduct constraints at the instance level. Thus they are difficult in handling complicated cluster structures and are limited in eliminating the impact of outliers in old embeddings, resulting in a risk of damaging the discriminative capability of new features. In this work, we propose a Neighborhood Consensus Contrastive Learning (NCCL) method. With no assumptions about the new training data, we estimate the sub-cluster structures of old embeddings. A new embedding is constrained with multiple old embeddings in both embedding space and discrimination space at the sub-class level. The effect of outliers diminished, as the multiple samples serve as ``mean teachers''. Besides, we propose a scheme to filter the old embeddings with low credibility, further improving the compatibility robustness. Our method ensures the compatibility without impairing the accuracy of the new model. It can even improve the new model's accuracy in most scenarios. Shengsen Wu, Yihang Lou, Minghua Deng, Ling-Yu Duan |
AAAI | 7 |
| 2022 | Fine-tuning Global Model via Data-Free Knowledge Distillation for Non-IID Federated LearningabstractFederated Learning (FL) is an emerging distributed learning paradigm under privacy constraint. Data heterogeneity is one of the main challenges in FL, which results in slow convergence and degraded performance. Most existing approaches only tackle the heterogeneity challenge by restricting the local model update in client, ignoring the performance drop caused by direct global model aggregation. Instead, we propose a data-free knowledge distillation method to fine-tune the global model in the server (FedFTG), which relieves the issue of direct model aggregation. Concretely, FedFTG explores the input space of local models through a generator, and uses it to transfer the knowledge from local models to the global model. Besides, we propose a hard sample mining scheme to achieve effective knowledge distillation throughout the training. In addition, we develop customized label sampling and class-level ensemble to derive maximum utilization of knowledge, which implicitly mitigates the distribution discrepancy across clients. Extensive experiments show that our FedFTG significantly outperforms the state-of-the-art (SOTA) FL algorithms and can serve as a strong plugin for enhancing FedAvg, FedProx, FedDyn, and SCAFFOLD. Lin Zhang 0014, Li Shen 0008, Liang Ding 0006, Dacheng Tao, Ling-Yu Duan |
CVPR | 5 |
| 2022 | mc-BEiT: Multi-choice Discretization for Image BERT Pre-training
Yixiao Ge, Ying Shan, Ling-Yu Duan |
ECCV (30) | 6 |
| 2022 | Uncertainty Modeling for Out-of-Distribution Generalization
Yongxing Dai, Yixiao Ge, Jun Liu 0036, Ying Shan, Ling-Yu Duan |
ICLR | 6 |
| 2022 | Collaborative Scalable Visual Compression for Human-Centered VideosabstractMachine intelligence systems have been increasingly widely deployed in real-world circumstances, while the conventional human-vision oriented video coding schemes are inefficient to be embedded in large-scale systems and further support a wide range of applications. There have been urgent demands for a new generation of compression framework to efficiently encodes visual data, where the compression and analytics for machine vision and human perception can be jointly optimized. To this end, we propose a novel visual compression framework to provide visual contents with different granularity for both human and machine vision tasks collaboratively. The proposed scalable compression framework maintains the critical semantic information in a basic layer, so that it is capable of supporting the accurate machine vision analysis under a tight bit-rate constraint. It is scalable to provide visual representations of different granularity to support various kinds of tasks, including video reconstruction that serves human vision examination. Experimental results on the human-centered videos have demonstrated the promising functionality of scalable visual coding with improved efficiency for high-performance machine analysis and human perception. Haofeng Huang, Wenhan Yang, Jiaying Liu 0001, Ling-Yu Duan |
ISCAS | 5 |
| 2022 | Nonlinear Multi-Model ReuseabstractThe goal of model reuse is to build a model in a new target domain by reusing some pre-trained source models. It can significantly reduce the training costs and the data required for training, and hence has various potential applications. Most of the existing model reuse approaches only reuse the output features or labels of the source model, and more information contained in the model are ignored. Besides, only a single model can be utilized in these approaches. A recently proposed multi-model reuse method is able to remedy these drawbacks by utilizing the hidden layer representations of multiple source models to help improve the representations in the target model, but it assumes that there are linear connections between the source and target models. This assumption is too restrictive and may be not valid in real-world applications. In this paper, we relax this assumption by introducing the manifold regularization scheme to exploit arbitrary nonlinear relationships between the source and target models. Effectiveness of our method is demonstrated empirically by the extensive experiments in the popular person re-identification task for smart city application. Yong Luo 0002, Ling-Yu Duan, Tongliang Liu, Yihang Lou, Yonggang Wen 0001 |
MMSP | 2 |
| 2022 | Disentangled Feature Learning Network and a Comprehensive Benchmark for Vehicle Re-IdentificationabstractVehicle Re-Identification (ReID) is of great significance for public security and intelligent transportation. Large and comprehensive datasets are crucial for the development of vehicle ReID in model training and evaluation. However, existing datasets in this field have limitations in many aspects, including the constrained capture conditions, limited variation of vehicle appearances, and small scale of training and test set, etc. Hence, a new, large, and challenging benchmark for vehicle ReID is urgently needed. In this paper, we propose a large vehicle ReID dataset, called VERI-Wild 2.0, containing 825,042 images. It is captured using a city-scale surveillance camera system, consisting of 274 cameras covering a very large area over 200$km^2$. Specifically, the samples in our dataset present very rich appearance diversities thanks to the long time span collecting settings, unconstrained capturing viewpoints, various illumination conditions, diversified background environments, and different weather conditions. Furthermore, to facilitate more practical benchmarking, we define a challenging and large test set containing about 400K vehicle images that do not have any camera overlap with the training set. VERI-Wild 2.0 is expected to be able to facilitate the design, adaptation, development, and evaluation of different types of learning models for vehicle ReID. Besides, we also design a new method for vehicle ReID. We observe that orientation is a crucial factor for feature matching in vehicle ReID. To match vehicle pairs captured from similar orientations, the learned features are expected to capture specific detailed differential information for discriminating the visually similar yet different vehicles. In contrast, features are desired to capture the orientation invariant common information when matching samples captured from different orientations. Thus a novel disentangled feature learning network (DFNet) is proposed. It explicitly considers the orientation information for vehicle ReID, and concurrently learns the orientation specific and orientation common features that thus can be adaptively exploited via an adaptive matching scheme when dealing with matching pairs from similar or different orientations. The comprehensive experimental results show the effectiveness of our proposed method. Jun Liu 0036, Yihang Lou, Ce Wang 0007, Ling-Yu Duan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Towards Low Light Enhancement With RAW ImagesabstractIn this paper, we make the first benchmark effort to elaborate on the superiority of using RAW images in the low light enhancement and develop a novel alternative route to utilize RAW images in a more flexible and practical way. Inspired by a full consideration on the typical image processing pipeline, we are inspired to develop a new evaluation framework, Factorized Enhancement Model (FEM), which decomposes the properties of RAW images into measurable factors and provides a tool for exploring how properties of RAW images affect the enhancement performance empirically. The empirical benchmark results show that the Linearity of data and Exposure Time recorded in meta-data play the most critical role, which brings distinct performance gains in various measures over the approaches taking the sRGB images as input. With the insights obtained from the benchmark results in mind, a RAW-guiding Exposure Enhancement Network (REENet) is developed, which makes trade-offs between the advantages and inaccessibility of RAW images in real applications in a way of using RAW images only in the training phase. REENet projects sRGB images into linear RAW domains to apply constraints with corresponding RAW images to reduce the difficulty of modeling training. After that, in the testing phase, our REENet does not rely on RAW images. Experimental results demonstrate not only the superiority of REENet to state-of-the-art sRGB-based methods and but also the effectiveness of the RAW guidance and all components. Haofeng Huang, Wenhan Yang, Yueyu Hu, Jiaying Liu 0001, Ling-Yu Duan |
IEEE Trans. Image Process. | 5 |
| 2022 | Intrinsic Performance Influence-based Participant Contribution Estimation for Horizontal Federated LearningabstractThe rapid development of modern artificial intelligence technique is mainly attributed to sufficient and high-quality data. However, in the data collection, personal privacy is at risk of being leaked. This issue can be addressed by federated learning, which is proposed to achieve efficient model training among multiple data providers without direct data access and aggregation. To encourage more parties owning high-quality data to participate in the federated learning, it is important to evaluate and reward the participant contribution in a reasonable, robust, and efficient manner. To achieve this goal, we propose a novel contribution estimation method: Intrinsic Performance Influence-based Contribution Estimation (IPICE). In particular, the class-level intrinsic performance influence is adopted as the contribution estimation criteria in IPICE, and a neural network is employed to exploit the non-linear relationship between the performance change and estimated contribution. Extensive experiments are conducted on various datasets, and the results demonstrate that IPICE is more accurate and stable than the counterpart in various data distribution settings. The computational complexity is significantly reduced in our IPICE, especially when a new party joins the federation. IPICE assigns small contributions to bad/garbage data and thus prevent them from participating and deteriorating the learning ecosystem. Lin Zhang 0014, Lixin Fan, Yong Luo 0002, Ling-Yu Duan |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2022 | Astute Video Transmission for Geographically Dispersed Devices in Visual IoT SystemsabstractVisual IoT (VIoT) is a promising IoT paradigm that visualizes sensing data from massive numbers of dispersed devices. A key objective in VIoT is to efficiently manage the devices to perform complex task-related visual data processing. Prior multimedia IoT systems have mainly focused on the delivery of captured video to remote servers, without considering the video tasks’ characteristics and the devices’ heterogeneous capabilities. In this work, we propose an astute video transmission framework for such a VIoT system composed of heterogeneous visual devices. First, we formulate the problem of joint video task allocation and heterogeneous device management by constructing a device hypergraph (DH) structure, which enables devices with different capabilities to perform complex video tasks cooperatively. Second, we model the video transmission within a VIoT system by applying fractal theory considering the NP-hardness of the optimization. In particular, we construct a comprehensive fractal submodular optimization framework through a DH and explore the inner submodular property to effectively leverage both video-task complexity and device heterogeneity. Third, we consider the geographically dispersed characteristic of massive numbers of VIoT devices and propose a multi-hop dispersed transmission mechanism for achieving globally cooperative optimality. The proposed architecture has been evaluated under diverse parameter settings. Numerical results are provided to validate the proposed algorithm in terms of delay, computational efficiency, and bandwidth utilization. Simulation results confirm the effectiveness and superiority of the proposed method. Wen Ji 0003, Ling-Yu Duan, Xi Huang 0002, Yueting Chai |
IEEE Trans. Mob. Comput. | 2 |
| 2022 | $A^3$-FKG: Attentive Attribute-Aware Fashion Knowledge Graph for Outfit Preference PredictionabstractWith the booming development of the online fashion industry, effective personalized recommender systems have become indispensable for the convenience they brought to the customers and the profits to the e-commercial platforms. Estimating the user’s preference towards the outfit is at the core of a personalized recommendation system. Existing works on fashion recommendation are largely centering on modelling the clothing compatibility without considering the user factor or characterizing the user’s preference over the single item. However, how to effectively model the outfits with either few or even none interactions, is yet under-explored. In this paper, we address the task of personalized outfit preference prediction via a novelAttentiveAttribute-AwareFashionKnowledgeGraph ($A^3$-FKG), which is incorporated to build the association between different outfits with both outfit- and item- level attributes. Additionally, a two-level attention mechanism is developed to capture the user’s preference: 1) User-specific relation-aware attention layer, which captures the user’s fine-grained preferences with different focus on relations for learning outfit representation; 2) Target-aware attention layer, which characterizes the user’s latent diverse interests from his/her behavior sequences for learning user representation. Extensive experiments conducted on a large-scale fashion outfit dataset demonstrate significant improvements over other methods, which verify the excellence of our proposed framework. Huijing Zhan, Jie Lin 0001, Kenan E. Ak, Boxin Shi, Ling-Yu Duan, Alex Chichung Kot |
IEEE Trans. Multim. | 5 |
| 2021 | Person30K: A Dual-Meta Generalization Network for Person Re-IdentificationabstractRecently, person re-identification (ReID) has vastly benefited from the surging waves of data-driven methods. However, these methods are still not reliable enough for real-world deployments, due to the insufficient generalization capability of the models learned on existing benchmarks that have limitations in multiple aspects, including limited data scale, capture condition variations, and appearance diversities. To this end, we collect a new dataset named Person30K with the following distinct features: 1) a very large scale containing 1.38 million images of 30K identities, 2) a large capture system containing 6,497 cameras deployed at 89 different sites, 3) abundant sample diversities including varied backgrounds and diverse person poses. Furthermore, we propose a domain generalization ReID method, dual-meta generalization network (DMG-Net), to exploit the merits of meta-learning in both the training procedure and the metric space learning. Concretely, we design a "learning then generalization evaluation" metatraining procedure and a meta-discrimination loss to enhance model generalization and discrimination capabilities. Comprehensive experiments validate the effectiveness of our DMG-Net. Jile Jiao, Ce Wang 0007, Jun Liu 0036, Yihang Lou, Xuetao Feng, Ling-Yu Duan |
CVPR | 7 |
| 2021 | Generalizable Person Re-Identification With Relevance-Aware Mixture of ExpertsabstractDomain generalizable (DG) person re-identification (ReID) is a challenging problem because we cannot access any unseen target domain data during training. Almost all the existing DG ReID methods follow the same pipeline where they use a hybrid dataset from multiple source domains for training, and then directly apply the trained model to the unseen target domains for testing. These methods often neglect individual source domains’ discriminative characteristics and their relevances w.r.t. the unseen target domains, though both of which can be leveraged to help the model’s generalization. To handle the above two issues, we propose a novel method called the relevance-aware mixture of experts (RaMoE), using an effective voting-based mixture mechanism to dynamically leverage source domains’ diverse characteristics to improve the model’s generalization. Specifically, we propose a decorrelation loss to make the source domain networks (experts) keep the diversity and discriminability of individual domains’ characteristics. Besides, we design a voting network to adaptively integrate all the experts’ features into the more generalizable aggregated features with domain relevance. Considering the target domains’ invisibility during training, we propose a novel learning-to-learn algorithm combined with our relation alignment loss to update the voting network. Extensive experiments demonstrate that our proposed RaMoE outperforms the state-of-the-art methods. Yongxing Dai, Jun Liu 0036, Zekun Tong, Ling-Yu Duan |
CVPR | 5 |
| 2021 | Single Image Reflection Removal With Absorption EffectabstractIn this paper, we consider the absorption effect for the problem of single image reflection removal. We show that the absorption effect can be numerically approximated by the average of refractive amplitude coefficient map. We then reformulate the image formation model and propose a two-step solution that explicitly takes the absorption effect into account. The first step estimates the absorption effect from a reflection-contaminated image, while the second step recovers the transmission image by taking a reflection-contaminated image and the estimated absorption effect as the input. Experimental results on four public datasets show that our two-step solution not only successfully removes reflection artifact, but also faithfully restores the intensity distortion caused by the absorption effect. Our ablation studies further demonstrate that our method achieves superior performance on the recovery of overall intensity and has good model generalization capacity. The code is available at https://github.com/q-zh/absorption. Boxin Shi, Jinnan Chen, Xudong Jiang 0001, Ling-Yu Duan, Alex Chichung Kot |
CVPR | 5 |
| 2021 | IDM: An Intermediate Domain Module for Domain Adaptive Person Re-IDabstractUnsupervised domain adaptive person re-identification (UDA re-ID) aims at transferring the labeled source domain’s knowledge to improve the model’s discriminability on the unlabeled target domain. From a novel perspective, we argue that the bridging between the source and target domains can be utilized to tackle the UDA re-ID task, and we focus on explicitly modeling appropriate intermediate domains to characterize this bridging. Specifically, we propose an Intermediate Domain Module (IDM) to generate intermediate domains’ representations on-the-fly by mixing the source and target domains’ hidden representations using two domain factors. Based on the "shortest geodesic path" definition, i.e., the intermediate domains along the shortest geodesic path between the two extreme domains can play a better bridging role, we propose two properties that these intermediate domains should satisfy. To ensure these two properties to better characterize appropriate intermediate domains, we enforce the bridge losses on intermediate domains’ prediction space and feature space, and enforce a diversity loss on the two domain factors. The bridge losses aim at guiding the distribution of appropriate intermediate domains to keep the right distance to the source and target domains. The diversity loss serves as a regularization to prevent the generated intermediate domains from being over-fitting to either of the source and target domains. Our proposed method outperforms the state-of-the-arts by a large margin in all the common UDA re-ID tasks, and the mAP gain is up to 7.7% on the challenging MSMT17 benchmark. Code is available at https://github.com/SikaStar/IDM. Yongxing Dai, Jun Liu 0036, Yifan Sun 0003, Zekun Tong, Chi Zhang 0026, Ling-Yu Duan |
ICCV | 6 |
| 2021 | Federated Learning for Non-IID Data via Unified Feature Learning and Optimization Objective AlignmentabstractFederated Learning (FL) aims to establish a shared model across decentralized clients under the privacy-preserving constraint. Despite certain success, it is still challenging for FL to deal with non-IID (non-independent and identical distribution) client data, which is a general scenario in real-world FL tasks. It has been demonstrated that the performance of FL will be reduced greatly under the non-IID scenario, since the discrepant data distributions will induce optimization inconsistency and feature divergence issues. Besides, naively minimizing an aggregate loss function in this scenario may have negative impacts on some clients and thus deteriorate their personal model performance. To address these issues, we propose a Unified Feature learning and Optimization objectives alignment method (FedUFO) for non-IID FL. In particular, an adversary module is proposed to reduce the divergence on feature representation among different clients, and two consensus losses are proposed to reduce the inconsistency on optimization objectives from two perspectives. Extensive experiments demonstrate that our FedUFO can outperform the state-of-the-art approaches, including the competitive one data-sharing method. Besides, FedUFO can enable more reasonable and balanced model performance among different clients. Lin Zhang 0014, Yong Luo 0002, Bo Du 0001, Ling-Yu Duan |
ICCV | 5 |
| 2021 | Person Retrieval with Conv-TransformerabstractPart-level features obtained by uniformly partitioning have attracted much attention in person re-identification. However, standard uniform part partitions may lead to within-part inconsistency across different samples, as shown in Figure 1. Attention mechanisms (e.g., refined part pooling) have been proposed to refine part division with enhanced consistency. Unfortunately, such mechanisms adopt single-headed convolutional structures, fail to fuse fine-grained part information. Besides, convolution-based schemes can maintain local positional information but cannot effectively pre-serve relative positions between parts. This paper proposes a new CNN-Transformer hyper architecture called the Person Retrieval with Conv-Transformer (PRCT). We integrate the multi-head self-attention and positional embedding module, which are the core ingredients of non-convolutional Transformer, with a CNN-based part-feature extractor to maintain more precise within-part consistency in feature aggregation. With PRCT, we can effectively eliminate the part mis-alignments when matching different samples. We conduct extensive evaluations on the MSMT17, DukeMTMC-ReID, and Market-1501 datasets and obtain state-of-the-art performance. Shengsen Wu, Ce Wang 0007, Ling-Yu Duan |
ICME | 4 |
| 2021 | Face Image Reflection Removal
Renjie Wan, Boxin Shi, Haoliang Li, Ling-Yu Duan, Alex Chichung Kot |
Int. J. Comput. Vis. | 4 |
| 2021 | Towards Large-Scale Object Instance Search: A Multi-Block N-Ary TrieabstractObject instance search is a challenging task with a wide range of applications, but the fast search with high accuracy has not been well solved yet. In this paper, we investigate the object instance search from a new perspective in terms of joint precision and computational cost optimization, and propose a novel index structure i.e., Multi-Block N-ary Trie (MBNT) to accelerate the exact r-neighbor search in the Hamming space. Comprehensive studies are first carried out to reveal the performance of exact and approximate nearest neighbor (NN) algorithms for object instance search. An interesting finding that the exact search is more promising for very compact binary codes (e.g., 64-bit and 128-bit) is analyzed. Along this vein, we introduce a Trie structure, i.e., MBNT, which is specifically designed for improving the exact NN search performance in the context of large-scale object instance search. To index the binary codes, a subset of continuous bits of a binary string, denoted as a block, is regarded as an atomic indexing element. As such, the problem of lookup misses can be addressed. Theoretical analyses are also provided to show that our MBNT scheme can incur less computational cost than other hash table-based methods. Extensive experimental results on the 100M dataset have demonstrated that our method achieves faster search speed while maintaining the promising search precision towards large-scale object instance search. Mangui Liang, Feng Gao 0014, Yi-Cheng Huang, Xinfeng Zhang 0001, Ling-Yu Duan |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | Digital Retina: A Way to Make the City Brain More Efficient by Visual CodingabstractThe ubiquitous camera networks in the city brain system grow at a rapid pace, creating massive amounts of images and videos at a range of spatial-temporal scales and thereby forming the “biggest” big data. However, the sensing system often lags behind the construction of the fast-growing city brain system, in the sense that such exponentially growing data far exceed today’s sensing capabilities. Therefore, critical issues arise regarding how to better leverage the existing city brain system and significantly improve the city-scale performance in intelligent applications. To tackle the unprecedented challenges, we articulate a vision towards a novel visual computing framework, termed asdigital retina, which aligns high-efficiency sensing models with the emerging Visual Coding for Machine (VCM) paradigm. In particular, digital retina may consist of video coding, feature coding, model coding, as well as their joint optimization. The digital retina is biologically-inspired, rooted on the widely accepted view that the retina encodes the visual information for human perception, and extracts features by the brain downstream areas to disentangle the visual objects. Within the digital retina framework, three streams, i.e., video stream, feature stream, and model stream, work collaboratively over the end-edge-cloud platform. In particular, the compressed video stream serves for human vision, the compact feature stream targets for machine vision, and the model stream incrementally updates deep learning models to improve the performance of human/machine vision tasks. We have developed a prototype to demonstrate the technical advantages of digital retina, and extensive experiments have been conducted to validate that it is able to effectively support the video big data analysis and retrieval in the intelligent city system. In particular, up to$7000\times $compression ratio could be realized for visual data compression while maintaining competitive performance with pristine signal in a series of visual analysis tasks. Wen Gao 0001, Siwei Ma 0001, Ling-Yu Duan, Yonghong Tian 0001, Peiyin Xing, Yaowei Wang 0001, Shanshe Wang, Huizhu Jia, Tiejun Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Hierarchical Connectivity-Centered Clustering for Unsupervised Domain Adaptation on Person Re-IdentificationabstractUnsupervised domain adaptation (UDA) on person Re-Identification (ReID) aims to transfer the knowledge from a labeled source domain to an unlabeled target domain. Recent works mainly optimize the ReID models with pseudo labels generated by unsupervised clustering on the target domain. However, the pseudo labels generated by the unsupervised clustering methods are often unreliable, due to the severe intra-person variations and complicated cluster structures in the practical application scenarios. In this work, to handle the complicated cluster structures, we propose a novel learnable Hierarchical Connectivity-Centered (HCC) clustering scheme by Graph Convolutional Networks (GCNs) to generate more reliable pseudo labels. Our HCC scheme learns the complicated cluster structure by hierarchically estimating the connectivity among samples from the vertex level to cluster level in a graph representation, and thereby progressively refines the pseudo labels. Additionally, to handle the intra-person variations in clustering, we propose a novel relation feature for HCC clustering, which exploits the identities from the source domain as references to represent target domain samples. Experiments demonstrate that our method is able to achieve state-of-the art performance on three challenging benchmarks. Ce Wang 0007, Yihang Lou, Jun Liu 0036, Ling-Yu Duan |
IEEE Trans. Image Process. | 5 |
| 2021 | Dual-Refinement: Joint Label and Feature Refinement for Unsupervised Domain Adaptive Person Re-IdentificationabstractUnsupervised domain adaptive (UDA) person re-identification (re-ID) is a challenging task due to the missing of labels for the target domain data. To handle this problem, some recent works adopt clustering algorithms to off-line generate pseudo labels, which can then be used as the supervision signal for on-line feature learning in the target domain. However, the off-line generated labels often contain lots of noise that significantly hinders the discriminability of the on-line learned features, and thus limits the final UDA re-ID performance. To this end, we propose a novel approach, called Dual-Refinement, that jointly refines pseudo labels at the off-line clustering phase and features at the on-line training phase, to alternatively boost the label purity and feature discriminability in the target domain for more reliable re-ID. Specifically, at the off-line phase, a new hierarchical clustering scheme is proposed, which selects representative prototypes for every coarse cluster. Thus, labels can be effectively refined by using the inherent hierarchical information of person images. Besides, at the on-line phase, we propose an instant memory spread-out (IM-spread-out) regularization, that takes advantage of the proposed instant memory bank to store sample features of the entire dataset and enable spread-out feature learning over the entire training data instantly. Our Dual-Refinement method reduces the influence of noisy labels and refines the learned features within the alternative training process. Experiments demonstrate that our method outperforms the state-of-the-art methods by a large margin. Yongxing Dai, Jun Liu 0036, Zekun Tong, Ling-Yu Duan |
IEEE Trans. Image Process. | 5 |
| 2021 | Towards Coding for Human and Machine Vision: Scalable Face Image CodingabstractThe past decades have witnessed the rapid development of image and video coding techniques in the era of big data. However, the signal fidelity-driven coding pipeline design limits the capability of the existing image/video coding frameworks to fulfill the needs of both machine and human vision. In this paper, we come up with a novel face image coding framework by leveraging both the compressive and the generative models, to support machine vision and human perception tasks jointly. Given an input image, the feature analysis is first applied, and then the generative model is employed to reconstruct image with compact structure and color features, where sparse edges are extracted to connect both kinds of vision and a key reference pixel selection method is proposed to determine the priorities of the reference color pixels for scalable coding. The compact edge map serves as the basic layer for machine vision tasks, and the reference pixels act as an enhanced layer to guarantee signal fidelity for human vision. By introducing advanced generative models, we train a decoding network to reconstruct images from compact structure and color representations, which is flexible to accept inputs in a scalable way and to control the imagery effect of the outputs between signal fidelity and visual realism. Experimental results and comprehensive performance analysis over the face image dataset demonstrate the superiority of our framework in both human vision tasks and machine vision tasks, which provide useful evidence on the emerging standardization efforts on MPEG VCM (Video Coding for Machine). Shuai Yang 0001, Yueyu Hu, Wenhan Yang, Ling-Yu Duan, Jiaying Liu 0001 |
IEEE Trans. Multim. | 4 |
| 2021 | Pose-Normalized and Appearance-Preserved Street-to-Shop Clothing Image Generation and Feature LearningabstractWe tackle the task of street-to-shop clothing image synthesis. Given a daily person image with a particular clothing item captured in the street scenario, we aim to synthesize the frontal facing view of that item in the shop scenario. This problem has the following challenges: 1) the distinct visual discrepancy between the street and shop scenario; 2) the severe shape deformation of clothing in the presence of an arbitrary human pose; 3) the preservation of fine-grained details during the process of clothing image generation. In this paper, we jointly solve these difficulties by proposing a Pose-Normalized and Appearance-Preserved Generative Adversarial Network (PNAP-GAN). More specifically, conditioned on the clothing-agnostic representation (i.e., clothing landmarks and semantic parsing map), we disentangle the shape and appearance synthesis in a coarse-to-fine framework. Moreover, a semantic embedding loss is introduced to guide the domain transfer in the semantic level (i.e., keeping the clothing attributes). With the synthesized frontal shop image, a pose-normalized representation in complementary to the domain-invariant feature learnt from the original street image are integrated to facilitate the problem of street-to-shop clothing retrieval. Extensive experiments conducted demonstrate the effectiveness of the proposed PNAP-GAN on generating high quality frontal-view images and the excellence of the learnt pose-normalized features on the retrieval task than existing methods. In addition, we demonstrate that the pose-normalized retrieval feature benefits the cross-scenario (i.e., street-to-shop) clothing image generation in a semantic-preserved manner. Huijing Zhan, Chenyu Yi, Boxin Shi, Jie Lin 0001, Ling-Yu Duan, Alex Chichung Kot |
IEEE Trans. Multim. | 5 |
| 2021 | Market2Dish: Health-aware Food RecommendationabstractWith the rising incidence of some diseases, such as obesity and diabetes, the healthy diet is arousing increasing attention. However, most existing food-related research efforts focus on recipe retrieval, user-preference-based food recommendation, cooking assistance, or the nutrition and calorie estimation of dishes, ignoring the personalized health-aware food recommendation. Therefore, in this work, we present a personalized health-aware food recommendation scheme, namely, Market2Dish, mapping the ingredients displayed in the market to the healthy dishes eaten at home. The proposed scheme comprises three components, namely, recipe retrieval, user health profiling, and health-aware food recommendation. In particular, recipe retrieval aims to acquire the ingredients available to the users and then retrieve recipe candidates from a large-scale recipe dataset. User health profiling is to characterize the health conditions of users by capturing the textual health-related information crawled from social networks. Specifically, to solve the issue that the health-related information is extremely sparse, we incorporate a word-class interaction mechanism into the proposed deep model to learn the fine-grained correlations between the textual tweets and pre-defined health concepts. For the health-aware food recommendation, we present a novel category-aware hierarchical memory network–based recommender to learn the health-aware user-recipe interactions for better food recommendation. Moreover, extensive experiments demonstrate the effectiveness of the health-aware food recommendation scheme. Wenjie Wang 0007, Ling-Yu Duan, Peiguang Jing, Xuemeng Song, Liqiang Nie |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Attribute-wise Explainable Fashion Compatibility ModelingabstractWith the boom of the fashion market and people’s daily needs for beauty, clothing matching has gained increased research attention. In a sense, tackling this problem lies in modeling the human notions of the compatibility between fashion items, i.e., Fashion Compatibility Modeling (FCM), which plays an important role in a wide bunch of commercial applications, including clothing recommendation and dressing assistant. Recent advances in multimedia processing have shown remarkable effectiveness in accurate compatibility evaluation. However, these studies work like a black box and cannot provide appropriate explanations, which are indeed of importance for gaining users’ trust and improving their experience. In fact, fashion experts usually explain the compatibility evaluation through the matching patterns between fashion attributes (e.g., a silk tank top cannot go with a knit dress). Inspired by this, we devise an attribute-wise explainable FCM solution, named ExFCM , which can simultaneously generate the item-level compatibility evaluation for input fashion items and the attribute-level explanations for the evaluation result. In particular, ExFCM consists of two key components: attribute-wise representation learning and attribute interaction modeling. The former works on learning the region-aware attribute representation for each item with the threshold global average pooling. Besides, the latter is responsible for compiling the attribute-level matching signals into the overall compatibility evaluation adaptively with the attentive interaction mechanism. Note that ExFCM is trained without any attribute-level compatibility annotations, which facilitates its practical applications. Extensive experiments on two real-world datasets validate that ExFCM can generate more accurate compatibility evaluations than the existing methods, together with reasonable explanations. Xin Yang 0008, Xuemeng Song, Fuli Feng, Haokun Wen, Ling-Yu Duan, Liqiang Nie |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2020 | Reflection Scene Separation From a Single ImageabstractFor images taken through glass, existing methods focus on the restoration of the background scene by regarding the reflection components as noise. However, the scene reflected by glass surface also contains important information to be recovered, especially for the surveillance or criminal investigations. In this paper, instead of removing reflection components from the mixture image, we aim at recovering reflection scenes from the mixture image. We first propose a strategy to obtain such ground truth and its corresponding input images. Then, we propose a two-stage framework to obtain the visible reflection scene from the mixture image. Specifically, we train the network with a shift-invariant loss which is robust to misalignment between the input and output images. The experimental results show that our proposed method achieves promising results. Renjie Wan, Boxin Shi, Haoliang Li, Ling-Yu Duan, Alex Chichung Kot |
CVPR | 4 |
| 2020 | What Does Plate Glass Reveal About Camera Calibration?abstractThis paper aims to calibrate the orientation of glass and the field of view of the camera from a single reflection-contaminated image. We show how a reflective amplitude coefficient map can be used as a calibration cue. Different from existing methods, the proposed solution is free from image contents. To reduce the impact of a noisy calibration cue estimated from a reflection-contaminated image, we propose two strategies: an optimization-based method that imposes part of though reliable entries on the map and a learning-based method that fully exploits all entries. We collect a dataset containing 320 samples as well as their camera parameters for evaluation. We demonstrate that our method not only facilitates a general single image camera calibration method that leverages image contents but also contributes to improving the performance of single image reflection removal. Furthermore, we show our byproduct output helps alleviate the ill-posed problem of estimating the panorama from a single image. Jinnan Chen, Zhan Lu, Boxin Shi, Xudong Jiang 0001, Kim-Hui Yap, Ling-Yu Duan, Alex Chichung Kot |
CVPR | 7 |
| 2020 | FHDe2Net: Full High Definition Demoireing Network
Ce Wang 0007, Boxin Shi, Ling-Yu Duan |
ECCV (22) | 4 |
| 2020 | HARD-Net: Hardness-AwaRe Discrimination Network for 3D Early Activity Prediction
Jun Liu 0036, Wei Zhang 0021, Ling-Yu Duan |
ECCV (11) | 4 |
| 2020 | Classes Matter: A Fine-Grained Adversarial Approach to Cross-Domain Semantic Segmentation
Wei Zhang 0031, Ling-Yu Duan, Tao Mei 0001 |
ECCV (14) | 4 |
| 2020 | Deep Product Quantization Module for Efficient Image RetrievalabstractProduct Quantization (PQ) is one of the most popular Approximate Nearest Neighbor (ANN) methods for large-scale image retrieval, bringing better performance than hashing based methods. In recent years, several works extend the hard quantization to soft quantization with specially designed deep neural architectures. We propose a simple but effective deep Product Quantization Module (PQM) to jointly learn discriminative codebook and precise hard assignment in an end-to-end manner. In this work, we use the straight-through estimator to make it feasible to directly optimize the discrete binary representations in deep neural networks with stochastic gradient descent. Different from previous deep vector quantization methods, PQM is a plug-and-play module which can be adaptive to various base networks in the scenarios of image search or compression. Besides, we propose a reconstruction loss to minimize the domain gap between the original embedding features and codebook. Experimental results show that PQM outperforms state-of-the-art deep supervised hashing and quantization methods on several image retrieval benchmarks. Meihan Liu, Yongxing Dai, Ling-Yu Duan |
ICASSP | 4 |
| 2020 | Data Representation in Hybrid Coding Framework for Feature Maps CompressionabstractRecently, a new paradigm of transmitting and compressing intermediate deep learning features (i.e., feature maps) for distributed visual analysis systems is emerging. As the fundamental infrastructure in such paradigm, research and standardization for feature maps coding has attracted more and more attention. In this paper, to improve the state-of-the-art hybrid coding framework which integrates the traditional video codecs to compress feature maps, we investigate the data representation procedure in such coding framework. Specifically, we proposed three modes in Repack module to help explore inter-channel redundancy, and we explore the fidelity maintenance ability of two modes in Pre-Quantization modules. It is worth mentioning that the proposed coding modes have been partially adopted in to the ongoing AVS (Audio Video Coding Standard Workgroup) - Visual Feature Coding Standard. Zhuo Chen 0006, Ling-Yu Duan, Shiqi Wang 0001, Weisi Lin, Alex Chichung Kot |
ICIP | 2 |
| 2020 | Extending Hashing Towards Fast Re-IdentificationabstractSearching accuracy and efficiency are two challenges in person and vehicle Re-identification (Re-ID), which one focuses on robust representations learning which usually generating high-dimensional features while the other has not been fully explored. Hashing is a suitable solution to make REID efficient. However, directly extending the existing hashing methods to fast Re-ID faces two challenges: one is the non-overlap between training and testing set which need more discriminative hash codes, the other is the large identities in Re-ID tasks which will lead to slow convergence and hard optimization. In this work, we propose an attention pooling operator to exploit both local and global visual attributes which can break limited discriminative power in hash methods. To further make training procedure converge faster and optimize the network more easily, we substitute non-differentiable l1-regularization with smooth l1-regularization. In experiments, our work outperforms state-of-the-art hashing and quantization methods on both person and vehicle Re-ID datasets. Besides, the results can serve as a strong baseline in the field of deep hashing for fast Re-ID.11Code is available at https://github.com/cynthia951031/Hashing ReID. Meihan Liu, Yongxing Dai, Shengsen Wu, Ling-Yu Duan |
ICIP | 5 |
| 2020 | Towards Coding For Human And Machine Vision: A Scalable Image Coding ApproachabstractThe past decades have witnessed the rapid development of image and video coding techniques in the era of big data. However, the signal fidelity-driven coding pipeline design limits the capability of the existing image/video coding frameworks to fulfill the needs of both machine and human vision. In this paper, we come up with a novel image coding framework by leveraging both the compressive and the generative models, to support machine vision and human perception tasks jointly. Given an input image, the feature analysis is first applied, and then the generative model is employed to perform image reconstruction with features and additional reference pixels, in which compact edge maps are extracted in this work to connect both kinds of vision in a scalable way. The compact edge map serves as the basic layer for machine vision tasks, and the reference pixels act as a sort of enhanced layer to guarantee signal fidelity for human vision. By introducing advanced generative models, we train a flexible network to reconstruct images from compact feature representations and the reference pixels. Experimental results demonstrate the superiority of our framework in both human visual quality and facial landmark detection, which provide useful evidence on the emerging standardization efforts on MPEG VCM (Video Coding for Machine). Our project website is available at https://williamyang1991.github.io/projects/VCM-Face/. Yueyu Hu, Shuai Yang 0001, Wenhan Yang, Ling-Yu Duan, Jiaying Liu 0001 |
ICME | 4 |
| 2020 | An Emerging Coding Paradigm Vcm: A Scalable Coding Approach Beyond Feature And SignalabstractIn this paper, we study a new problem arising from the emerging MPEG standardization effort Video Coding for Machine (VCM)1, which aims to bridge the gap between visual feature compression and classical video coding. VCM is committed to address the requirement of compact signal representation for both machine and human vision in a more or less scalable way. To this end, we make endeavors in leveraging the strength of predictive and generative models to support advanced compression techniques for both machine and human vision tasks simultaneously, in which visual features serve as a bridge to connect signal-level and task-level compact representations in a scalable manner. Specifically, we employ a conditional deep generation network to reconstruct video frames with the guidance of learned motion pattern. By learning to extract sparse motion pattern via a predictive model, the network elegantly leverages the feature representation to generate the appearance of to-be-coded frames via a generative model, relying on the appearance of the coded key frames. Meanwhile, the sparse motion pattern is compact and highly effective for high-level vision tasks, e.g. action recognition. Experimental results demonstrate that our method yields much better reconstruction quality compared with the traditional video codecs (0.0063 gain in SSIM), as well as state-of-the-art action recognition performance over highly compressed videos (9.4% gain in recognition accuracy), which showcases a promising paradigm of coding signal for both human and machine vision. Sifeng Xia, Kunchangtai Liang, Wenhan Yang, Ling-Yu Duan, Jiaying Liu 0001 |
ICME | 4 |
| 2020 | Disentangled Feature Learning Network for Vehicle Re-IdentificationabstractVehicle Re-Identification (ReID) has attracted lots of research efforts due to its great significance to the public security. In vehicle ReID, we aim to learn features that are powerful in discriminating subtle differences between vehicles which are visually similar, and also robust against different orientations of the same vehicle. However, these two characteristics are hard to be encapsulated into a single feature representation simultaneously with unified supervision. Here we propose a Disentangled Feature Learning Network (DFLNet) to learn orientation specific and common features concurrently, which are discriminative at details and invariant to orientations, respectively. Moreover, to effectively use these two types of features for ReID, we further design a feature metric alignment scheme to ensure the consistency of the metric scales. The experiments show the effectiveness of our method that achieves state-of-the-art performance on three challenging datasets. Yihang Lou, Yongxing Dai, Jun Liu 0036, Ziqian Chen, Ling-Yu Duan |
IJCAI | 6 |
| 2020 | Pose-native Network Architecture Search for Multi-person Human Pose EstimationabstractMulti-person pose estimation has achieved great progress in recent years, even though, the precise prediction for occluded and invisible hard keypoints remains challenging. Most of the human pose estimation networks are equipped with an image classification-based pose encoder for feature extraction and a handcrafted pose decoder for high-resolution representations. However, the pose encoder might be sub-optimal because of the gap between image classification and pose estimation. The widely used multi-scale feature fusion in pose decoder is still coarse and cannot provide sufficient high-resolution details for hard keypoints. Neural Architecture Search (NAS) has shown great potential in many visual tasks to automatically search efficient networks. In this work, we present the Pose-native Network Architecture Search (PoseNAS) to simultaneously design a better pose encoder and pose decoder for pose estimation. Specifically, we directly search a data-oriented pose encoder with stacked searchable cells, which can provide an optimum feature extractor for the pose specific task. In the pose decoder, we exploit scale-adaptive fusion cells to promote rich information exchange across the multi-scale feature maps. Meanwhile, the pose decoder adopts a Fusion-and-Enhancement manner to progressively boost the high-resolution representations that are non-trivial for the precious prediction of hard keypoints. With the exquisitely designed search space and search strategy, PoseNAS can simultaneously search all modules in an end-to-end manner. PoseNAS achieves state-of-the-art performance on three public datasets, MPII, COCO, and PoseTrack, with small-scale parameters compared with the existing methods. Our best model obtains 76.7% mAP and 75.9% mAP on the COCO validation set and test set with only 33.6M parameters. Code and implementation are available at https://github.com/for-code0216/PoseNAS. Qian Bao, Wu Liu 0005, Ling-Yu Duan, Tao Mei 0001 |
ACM Multimedia | 4 |
| 2020 | Network Update Compression for Federated LearningabstractIn federated learning setting, models are trained in a variety of edge-devices with locally generated data and each round only updates in the current model rather than the model itself are sent to the server where they are aggregated to compose an improved model. These edge devices, however, reside in highly uneven nature of network with higher latency and lower-throughput connections and are intermittently available for training. In addition, a network connection has an asymmetric nature of downlink and uplink. All these contribute to a major challenge while synchronizing these updates to the server.In this work, we proposed an efficient coding solution to significantly reduce uplink communication cost by reducing the total number of parameters required for updates. This was achieved by applying Gaussian Mixture Model (GMM) to localize Karhunen-Loève Transform (KLT) on inter-model subspace and representing it with two low-rank matrices. Experiments on convolutional neural network (CNN) models showed the proposed model can significantly reduce the uplink communication cost in federated learning while preserving reasonable accuracy. Birendra Kathariya, Li Li 0040, Zhu Li 0001, Ling-Yu Duan, Shan Liu 0001 |
VCIP | 4 |
| 2020 | Deep Variational and Structural HashingabstractIn this paper, we propose a deep variational and structural hashing (DVStH) method to learn compact binary codes for multimedia retrieval. Unlike most existing deep hashing methods which use a series of convolution and fully-connected layers to learn binary features, we develop a probabilistic framework to infer latent feature representation inside the network. Then, we design a struct layer rather than a bottleneck hash layer, to obtain binary codes through a simple encoding procedure. By doing these, we are able to obtain binary codes discriminatively and generatively. To make it applicable to cross-modal scalable multimedia retrieval, we extend our method to a cross-modal deep variational and structural hashing (CM-DVStH). We design a deep fusion network with a struct layer to maximize the correlation between image-text input pairs during the training stage so that a unified binary vector can be obtained. We then design modality-specific hashing networks to handle the out-of-sample extension scenario. Specifically, we train a network for each modality which outputs a latent representation that is as close as possible to the binary codes which are inferred from the fusion network. Experimental results on five benchmark datasets are presented to show the efficacy of the proposed approach. Venice Erin Liong, Jiwen Lu, Ling-Yu Duan, Yap-Peng Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Feature Boosting Network For 3D Pose EstimationabstractIn this paper, a feature boosting network is proposed for estimating 3D hand pose and 3D body pose from a single RGB image. In this method, the features learned by the convolutional layers are boosted with a new long short-term dependence-aware (LSTD) module, which enables the intermediate convolutional feature maps to perceive the graphical long short-term dependency among different hand (or body) parts using the designed Graphical ConvLSTM. Learning a set of features that are reliable and discriminatively representative of the pose of a hand (or body) part is difficult due to the ambiguities, texture and illumination variation, and self-occlusion in the real application of 3D pose estimation. To improve the reliability of the features for representing each body part and enhance the LSTD module, we further introduce a context consistency gate (CCG) in this paper, with which the convolutional feature maps are modulated according to their consistency with the context representations. We evaluate the proposed method on challenging benchmark datasets for 3D hand pose estimation and 3D full body pose estimation. Experimental results show the effectiveness of our method that achieves state-of-the-art performance on both of the tasks. Jun Liu 0036, Henghui Ding, Amir Shahroudy, Ling-Yu Duan, Xudong Jiang 0001, Gang Wang 0012, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity UnderstandingabstractResearch on depth-based human activity analysis achieved outstanding performance and demonstrated the effectiveness of 3D representation for action recognition. The existing depth-based and RGB+D-based action recognition benchmarks have a number of limitations, including the lack of large-scale training samples, realistic number of distinct class categories, diversity in camera views, varied environmental conditions, and variety of human subjects. In this work, we introduce a large-scale dataset for RGB+D human action recognition, which is collected from 106 distinct subjects and contains more than 114 thousand video samples and 8 million frames. This dataset contains 120 different action classes including daily, mutual, and health-related activities. We evaluate the performance of a series of existing 3D activity analysis methods on this dataset, and show the advantage of applying deep learning methods for 3D-based human action recognition. Furthermore, we investigate a novel one-shot 3D activity recognition problem on our dataset, and a simple yet effective Action-Part Semantic Relevance-aware (APSR) framework is proposed for this task, which yields promising results for recognition of the novel action classes. We believe the introduction of this large-scale dataset will enable the community to apply, adapt, and develop various data-hungry learning techniques for depth-based and RGB+D-based human activity understanding. Jun Liu 0036, Amir Shahroudy, Mauricio Perez, Gang Wang 0012, Ling-Yu Duan, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2020 | Skeleton-Based Online Action Prediction Using Scale Selection NetworkabstractAction prediction is to recognize the class label of an ongoing activity when only a part of it is observed. In this paper, we focus on online action prediction in streaming 3D skeleton sequences. A dilated convolutional network is introduced to model the motion dynamics in temporal dimension via a sliding window over the temporal axis. Since there are significant temporal scale variations in the observed part of the ongoing action at different time steps, a novel window scale selection method is proposed to make our network focus on the performed part of the ongoing action and try to suppress the possible incoming interference from the previous actions at each step. An activation sharing scheme is also proposed to handle the overlapping computations among the adjacent time steps, which enables our framework to run more efficiently. Moreover, to enhance the performance of our framework for action prediction with the skeletal input data, a hierarchy of dilated tree convolutions are also designed to learn the multi-level structured semantic representations over the skeleton joints at each frame. Our proposed approach is evaluated on four challenging datasets. The extensive experiments demonstrate the effectiveness of our method for skeleton-based online action prediction. Jun Liu 0036, Amir Shahroudy, Gang Wang 0012, Ling-Yu Duan, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | CoRRN: Cooperative Reflection Removal NetworkabstractRemoving the undesired reflections from images taken through the glass is of broad application to various computer vision tasks. Non-learning based methods utilize different handcrafted priors such as the separable sparse gradients caused by different levels of blurs, which often fail due to their limited description capability to the properties of real-world reflections. In this paper, we propose a network with the feature-sharing strategy to tackle this problem in a cooperative and unified framework, by integrating image context information and the multi-scale gradient information. To remove the strong reflections existed in some local regions, we propose a statistic loss by considering the gradient level statistics between the background and reflections. Our network is trained on a new dataset with 3250 reflection images taken under diverse real-world scenes. Experiments on a public benchmark dataset show that the proposed method performs favorably against state-of-the-art methods. Renjie Wan, Boxin Shi, Haoliang Li, Ling-Yu Duan, Ah-Hwee Tan, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Toward Intelligent Sensing: Intermediate Deep Feature CompressionabstractThe recent advances of hardware technology have made the intelligent analysis equipped at the front-end with deep learning more prevailing and practical. To better enable the intelligent sensing at the front-end, instead of compressing and transmitting visual signals or the ultimately utilized top-layer deep learning features, we propose to compactly represent and convey the intermediate-layer deep learning features with high generalization capability, to facilitate the collaborating approach between front and cloud ends. This strategy enables a good balance among the computational load, transmission load and the generalization ability for cloud servers when deploying the deep neural networks for large scale cloud based visual analysis. Moreover, the presented strategy also makes the standardization of deep feature coding more feasible and promising, as a series of tasks can simultaneously benefit from the transmitted intermediate layer features. We also present the results for evaluations of both lossless and lossy deep feature compression, which provide meaningful investigations and baselines for future research and standardization activities. Zhuo Chen 0006, Kui Fan, Shiqi Wang 0001, Ling-Yu Duan, Weisi Lin, Alex Chichung Kot |
IEEE Trans. Image Process. | 4 |
| 2020 | Video Coding for Machines: A Paradigm of Collaborative Compression and Intelligent AnalyticsabstractVideo coding, which targets to compress and reconstruct the whole frame, and feature compression, which only preserves and transmits the most critical information, stand at two ends of the scale. That is, one is with compactness and efficiency to serve for machine vision, and the other is with full fidelity, bowing to human perception. The recent endeavors in imminent trends of video compression, e.g. deep learning based coding tools and end-to-end image/video coding, and MPEG-7 compact feature descriptor standards, i.e. Compact Descriptors for Visual Search and Compact Descriptors for Video Analysis, promote the sustainable and fast development in their own directions, respectively. In this paper, thanks to booming AI technology, e.g. prediction and generation models, we carry out exploration in the new area, Video Coding for Machines (VCM), arising from the emerging MPEG standardization efforts1. Towards collaborative compression and intelligent analytics, VCM attempts to bridge the gap between feature coding for machine vision and video coding for human vision. Aligning with the rising Analyze then Compress instance Digital Retina, the definition, formulation, and paradigm of VCM are given first. Meanwhile, we systematically review state-of-the-art techniques in video compression and feature compression from the unique perspective of MPEG standardization, which provides the academic and industrial evidence to realize the collaborative compression of video and feature streams in a broad range of AI applications. Finally, we come up with potential VCM solutions, and the preliminary results have demonstrated the performance and efficiency gains. Further direction is discussed as well. Ling-Yu Duan, Jiaying Liu 0001, Wenhan Yang, Tiejun Huang 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Iterative Local-Global Collaboration Learning Towards One-Shot Video Person Re-IdentificationabstractVideo person re-identification (video Re-ID) plays an important role in surveillance video analysis and has gained increasing attention recently. However, existing supervised methods require vast labeled identities across cameras, resulting in poor scalability in practical applications. Although some unsupervised approaches have been exploited for video Re-ID, they are still in their infancy due to the complex nature of learning discriminative features on unlabelled data. In this paper, we focus on one-shot video Re-ID and present an iterative local-global collaboration learning approach to learning robust and discriminative person representations. Specifically, it jointly considers the global video information and local frame sequence information to better capture the diverse appearance of the person for feature learning and pseudo-label estimation. Moreover, as the cross-entropy loss may induce the model to focus on identity-irrelevant factors, we introduce the variational information bottleneck as a regularization term to train the model together. It can help filter undesirable information and characterize subtle differences among persons. Since accuracy cannot always be guaranteed for pseudo-labels, we adopt a dynamic selection strategy to select part of pseudo-labeled data with higher confidence to update the training set and re-train the learning model. During training, our method iteratively executes the feature learning, pseudo-label estimation, and dynamic sample selection until all the unlabeled data have been seen. Extensive experiments on two public datasets, i.e., DukeMTMC-VideoReID and MARS, have verified the superiority of our model to several cutting-edge competitors. Meng Liu 0006, Leigang Qu, Liqiang Nie, Maofu Liu, Ling-Yu Duan, Baoquan Chen |
IEEE Trans. Image Process. | 5 |
| 2020 | Towards Efficient Front-End Visual Sensing for Digital Retina: A Model-Centric ParadigmabstractThe digital retina excels at providing enhanced visual sensing and analysis capability for city brain in smart cities, and can feasibly convert the visual data from visual sensors into semantic features. With the deployment of deep learning or handcrafted models, these features are extracted on front-end devices, then delivered to back-end servers for advanced analysis. In this scenario, we propose a model generation, utilization and communication paradigm, aiming at strong front-end sensing capabilities for establishing better artificial visual systems in smart cities. In particular, we propose an integrated multiple deep learning models reuse and prediction strategy, which dramatically increases the feasibility of the digital retina in large-scale visual data analysis in smart cities. The proposed multi-model reuse scheme aims to reuse the knowledge from models cached and transmitted in digital retina to obtain more discriminative capability. To efficiently deliver these newly generated models, a model prediction scheme is further proposed by encoding and reconstructing model differences. Extensive experiments have been conducted to demonstrate the effectiveness of proposed model-centric paradigm. Yihang Lou, Ling-Yu Duan, Yong Luo 0002, Ziqian Chen, Tongliang Liu, Shiqi Wang 0001, Wen Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2019 | Exploring Object Relation in Mean Teacher for Cross-Domain DetectionabstractRendering synthetic data (e.g., 3D CAD-rendered images) to generate annotations for learning deep models in vision tasks has attracted increasing attention in recent years. However, simply applying the models learnt on synthetic images may lead to high generalization error on real images due to domain shift. To address this issue, recent progress in cross-domain recognition has featured the Mean Teacher, which directly simulates unsupervised domain adaptation as semi-supervised learning. The domain gap is thus naturally bridged with consistency regularization in a teacher-student scheme. In this work, we advance this Mean Teacher paradigm to be applicable for cross-domain detection. Specifically, we present Mean Teacher with Object Relations (MTOR) that novelly remolds Mean Teacher under the backbone of Faster R-CNN by integrating the object relations into the measure of consistency cost between teacher and student modules. Technically, MTOR firstly learns relational graphs that capture similarities between pairs of regions for teacher and student respectively. The whole architecture is then optimized with three consistency regularizations: 1) region-level consistency to align the region-level predictions between teacher and student, 2) inter-graph consistency for matching the graph structures between teacher and student, and 3) intra-graph consistency to enhance the similarity between regions of same class within the graph of student. Extensive experiments are conducted on the transfers across Cityscapes, Foggy Cityscapes, and SIM10k, and superior results are reported when comparing to state-of-the-art approaches. More remarkably, we obtain a new record of single model: 22.8% of mAP on Syn2Real detection dataset. Yingwei Pan, Chong-Wah Ngo, Xinmei Tian 0001, Ling-Yu Duan, Ting Yao 0003 |
CVPR | 5 |
| 2019 | Towards Accurate One-Stage Object Detection With AP-LossabstractOne-stage object detectors are trained by optimizing classification-loss and localization-loss simultaneously, with the former suffering much from extreme foreground-background class imbalance issue due to the large number of anchors. This paper alleviates this issue by proposing a novel framework to replace the classification task in one-stage detectors with a ranking task, and adopting the Average-Precision loss (AP-loss) for the ranking problem. Due to its non-differentiability and non-convexity, the AP-loss cannot be optimized directly. For this purpose, we develop a novel optimization algorithm, which seamlessly combines the error-driven update scheme in perceptron learning and backpropagation algorithm in deep networks. We verify good convergence property of the proposed algorithm theoretically and empirically. Experimental results demonstrate notable performance improvement in state-of-the-art one-stage detectors based on AP-loss over different kinds of classification-losses on various benchmarks, without changing the network architectures. Kean Chen, Weiyao Lin, John See, Ling-Yu Duan, Zhibo Chen 0006, Changwei He, Junni Zou |
CVPR | 6 |
| 2019 | VERI-Wild: A Large Dataset and a New Method for Vehicle Re-Identification in the WildabstractVehicle Re-identification (ReID) is of great significance to the intelligent transportation and public security. However, many challenging issues of Vehicle ReID in real-world scenarios have not been fully investigated, e.g., the high viewpoint variations, extreme illumination conditions, complex backgrounds, and different camera sources. To promote the research of vehicle ReID in the wild, we collect a new dataset called VERI-Wild with the following distinct features: 1) The vehicle images are captured by a large surveillance system containing 174 cameras covering a large urban district (more than 200km^2) The camera network continuously captures vehicles for 24 hours in each day and lasts for 1 month. 3) It is the first vehicle ReID dataset that is collected from unconstrained conditionsns. It is also a large dataset containing more than 400 thousand images of 40 thousand vehicle IDs. In this paper, we also propose a new method for vehicle ReID, in which, the ReID model is coupled into a Feature Distance Adversarial Network (FDA-Net), and a novel feature distance adversary scheme is designed to generate hard negative samples in feature space to facilitate ReID model training. The comprehensive results show the effectiveness of our method on the proposed dataset and the other two existing datasets. Yihang Lou, Jun Liu 0036, Shiqi Wang 0001, Ling-Yu Duan |
CVPR | 5 |
| 2019 | Separable KLT for Intra Coding in Versatile Video Coding (VVC)abstractAfter the works on the state-of-the-art High Efficiency Video Coding (HEVC) standard, the standard organizations continued to study the potential video coding technologies for the next generation of video coding standard, named Versatile Video Coding (VVC). Transform is a key technique for compression efficiency, and core experiment 6 (CE6) is carried out to explore the transform related coding tools. In this paper, we propose a novel separable transform based on Karhunen-Loève Transform (KLT) to eliminate the horizontal and vertical correlations in the residual samples of intra coding. In the proposed method, the weaknesses of the traditional KLT are addressed. The separable KLT is developed as an alternative transform type in addition to DCT-II, and the transform matrices from 4×4 to 64×64 are trained from intra residual samples. Experimental results show the proposed method can achieve 2.7% bitrate saving averagely on top of the reference software of VVC (VTM-1.1), and the consistent performance improvement on test set also validates the strong generalization capacity of the proposed separable KLT. Kui Fan, Ronggang Wang, Weisi Lin, Jong-Uk Hou, Ling-Yu Duan, Ge Li 0002, Wen Gao 0001 |
DCC | 5 |
| 2019 | Mop Moiré Patterns Using MopNetabstractMoiré pattern is a common image quality degradation caused by frequency aliasing between monitors and cameras when taking screen-shot photos. The complex frequency distribution, imbalanced magnitude in colour channels, and diverse appearance attributes of moiré pattern make its removal a challenging problem. In this paper, we propose a Moiré pattern Removal Neural Network (MopNet) to solve this problem. All core components of MopNet are specially designed for unique properties of moire patterns, including the multi-scale feature aggregation addressing complex frequency, the channel-wise target edge predictor to exploit imbalanced magnitude among colour channels, and the attribute-aware classifier to characterize the diverse appearance for better modelling Moiré patterns. Quantitative and qualitative experimental comparison validate the state-of-the-art performance of MopNet. Ce Wang 0007, Boxin Shi, Ling-Yu Duan |
ICCV | 4 |
| 2019 | Sampling Wisely: Deep Image Embedding by Top-K Precision OptimizationabstractDeep image embedding aims at learning a convolutional neural network (CNN) based mapping function that maps an image to a feature vector. The embedding quality is usually evaluated by the performance in image search tasks. Since very few users bother to open the second page search results, top-k precision mostly dominates the user experience and thus is one of the crucial evaluation metrics for the embedding quality. Despite being extensively studied, existing algorithms are usually based on heuristic observation without theoretical guarantee. Consequently, gradient descent direction on the training loss is mostly inconsistent with the direction of optimizing the concerned evaluation metric. This inconsistency certainly misleads the training direction and degrades the performance. In contrast to existing works, in this paper, we propose a novel deep image embedding algorithm with end-to-end optimization to top-k precision, the evaluation metric that is closely related to user experience. Specially, our loss function is constructed with wisely selected ``misplaced" images along the top k nearest neighbor decision boundary, so that the gradient descent update directly promotes the concerned metric, top-k precision. Further more, our theoretical analysis on the upper bounding and consistency properties of the proposed loss supports that minimizing our proposed loss is equivalent to maximizing top-k precision. Experiments show that our proposed algorithm outperforms all compared state-of-the-art deep image embedding algorithms on three benchmark datasets. Chaofan Xu, Wei Zhang 0031, Ling-Yu Duan, Tao Mei 0001 |
ICCV | 4 |
| 2019 | Learning to Jointly Generate and Separate ReflectionsabstractExisting learning-based single image reflection removal methods using paired training data have fundamental limitations about the generalization capability on real-world reflections due to the limited variations in training pairs. In this work, we propose to jointly generate and separate reflections within a weakly-supervised learning framework, aiming to model the reflection image formation more comprehensively with abundant unpaired supervision. By imposing the adversarial losses and combinable mapping mechanism in a multi-task structure, the proposed framework elegantly integrates the two separate stages of reflection generation and separation into a unified model. The gradient constraint is incorporated into the concurrent training process of the multi-task learning as well. In particular, we built up an unpaired reflection dataset with 4,027 images, which is useful for facilitating the weakly-supervised learning of reflection removal model. Extensive experiments on a public benchmark dataset show that our framework performs favorably against state-of-the-art methods and consistently produces visually appealing results. Daiqian Ma, Renjie Wan, Boxin Shi, Alex Chichung Kot, Ling-Yu Duan |
ICCV | 5 |
| 2019 | SPLINE-Net: Sparse Photometric Stereo Through Lighting Interpolation and Normal Estimation NetworksabstractThis paper solves the Sparse Photometric stereo through Lighting Interpolation and Normal Estimation using a generative Network (SPLINE-Net). SPLINE-Net contains a lighting interpolation network to generate dense lighting observations given a sparse set of lights as inputs followed by a normal estimation network to estimate surface normals. Both networks are jointly constrained by the proposed symmetric and asymmetric loss functions to enforce isotropic constrain and perform outlier rejection of global illumination effects. SPLINE-Net is verified to outperform existing methods for photometric stereo of general BRDFs by using only ten images of different lights instead of using nearly one hundred images. Yiming Jia, Boxin Shi, Xudong Jiang 0001, Ling-Yu Duan, Alex Chichung Kot |
ICCV | 5 |
| 2019 | Fashion Recommendation on Street ImagesabstractLearning the compatibility relationship is of vital importance to a fashion recommendation system, while existing works achieve this merely on product images but not on street images in the complex daily life scenario. In this paper, we propose a novel fashion recommendation system: Given a query item of interest in the street scenario, the system can return the compatible items. More specifically, a two-stage curriculum learning scheme is developed to transfer the semantics from the product to street outfit images. We also propose a domain-specific missing item imputation method based on style and color similarity to handle the incomplete outfits. To support the training of deep recommendation model, we collect a large dataset with street outfit images. The experiments on the dataset demonstrate the advantages of the proposed method over the state-of-the-art approaches on both the street images and the product images. Huijing Zhan, Boxin Shi, Ling-Yu Duan, Alex Chichung Kot |
ICIP | 5 |
| 2019 | Denoising Adversarial Networks for Rain Removal and Reflection RemovalabstractThis paper presents a novel adversarial scheme to perform image denoising for the tasks of rain streak removal and reflection removal. Similar to several previous works, the proposed method first estimates a prior image and then uses it to guide the inference of noise-free image. The novelty of our approach is to jointly learn the gradient and noise-free image based on an adversarial scheme. More specifically, we use the gradient map as the prior image. The inferred noise-free image guided by an estimated gradient is regarded as a negative sample, while the noise-free image guided by the ground truth of a gradient is taken as a positive sample. With the anchor defined by the ground truth of noise-free image, we play a min-max game to jointly train two optimizers for the estimation of the gradient and the inference of noise-free images. We show that both prior image and noise-free image can be accurately obtained under this adversarial scheme. Our state-of-the-art performance achieved on two public benchmark datasets validate the effectiveness of our approach. Boxin Shi, Xudong Jiang 0001, Ling-Yu Duan, Alex Chichung Kot |
ICIP | 4 |
| 2019 | Incorporating Category Taxonomy in Deep Reinforcement Learning Based Image HashingabstractImage hashing is critical for large-scale image analytic-based applications, such as image retrieval. Although there have been dozens of hashing approaches, few of them take the hierarchical structure of the image categories into consideration. In this paper, we propose to incorporate the category taxonomy information in a deep reinforcement learning (DRL) model for image hashing. In particular, we learn an agent to predict the hashing codes sequentially under the DRL theme. Each coordinate of the hashing function can take the errors incurred by previous ones into consideration and hence more reliable hashing codes can be obtained than learning them independently. Besides, we design a novel level-specific reward function to gradually refine the hashing function according to the taxonomy information. Extensive experiments on two popular datasets demonstrate effectiveness of the proposed method. Qiang Fu 0006, Linsen Dong, Yong Luo 0002, Yonggang Wen 0001, Ying Li 0012, Ling-Yu Duan |
ICME | 7 |
| 2019 | Towards Digital Retina in Smart Cities: A Model Generation, Utilization and Communication ParadigmabstractThe digital retina in smart cities is to select what the City Eye tells the City Brain, and convert the acquired visual data from front-end visual sensors to features in an intelligent sensing manner. By deploying deep learning and/or handcrafted models in front-end devices, the compact features can be extracted and subsequently delivered to back-end cloud for search and advanced analytics. In this context, we propose a model generation, utilization, and communication paradigm, aiming to address a set of unique challenges for better artificial intelligence services in smart cities. In particular, we present an integrated multiple deep learning models reuse and prediction strategy, which greatly increases the feasibility of the digital retina in processing and analyzing the large-scale visual data in smart cities. The promise of the proposed paradigm is demonstrated through a set of experiments. Yihang Lou, Ling-Yu Duan, Yong Luo 0002, Ziqian Chen, Tongliang Liu, Shiqi Wang 0001, Wen Gao 0001 |
ICME | 2 |
| 2019 | Learning to Remove Reflections for Text ImagesabstractText images taken behind a piece of glass in the wild are largely contaminated by reflections. Directly applying existing reflection removal methods on text images with reflections cannot recover clear and correct text contents due to the ignorance of special characteristics of texts. This paper proposes a stacked framework to solve the text image reflection removal problem by specifically considering the regional properties of reflection and embedding the specific text priors into the estimation process in a unified manner. Experiment results on a newly collected dataset demonstrate that the proposed method outperforms state-of-the-art methods in recovering visually pleasant reflection-free images and recognizable text features. Ce Wang 0007, Renjie Wan, Feng Gao 0014, Boxin Shi, Ling-Yu Duan |
ICME | 5 |
| 2019 | From Market to Dish: Multi-ingredient Image Recognition for Personalized Recipe RecommendationabstractRecognition of food ingredients enables applications on recipe recommendation for developing a healthier eating habit. Existing ingredients recognition methods largely rely on ideal images captured in a controlled environment, while ingredients are usually displayed unorderly in a complex environment in the market. We propose the multi-ingredient recognition problem in the market and develop a Spatial Regularization Network (SRN) based method to solve it by using a newly collected multiple vegetable image dataset captured in the market. We further use the recognition result to develop a recipe recommendation system to satisfy the daily nutrition requirements and individual preference of each user. Experiments show that our multi-ingredient recognition outperforms previous methods over 14% in mAP and recommendation model shows an improvement of over 23% in HR@10. Lin Zhang 0014, Jianbo Zhao 0002, Si Li 0001, Boxin Shi, Ling-Yu Duan |
ICME | 5 |
| 2019 | Few-Shot and Many-Shot Fusion Learning in Mobile Visual Food RecognitionabstractMobile visual food recognition is emerging as an important application in food logging and dietary monitoring in recent years. Existing food recognition methods use conventional many-shot learning to train a large backbone network, which refers to the use of sufficient number of training data to train the network. However, these methods firstly do not consider the cases where certain food categories have limited training data. Therefore, they cannot use the conventional training using many-shot learning. Further, existing solutions focus on improving the food recognition performance by implementing state-of-the-art large full networks, and do not pay much attention to reduce the size and computational cost of the network. As a result, they are not amenable for deployment on mobile devices. In this paper, we address these issues by proposing a new few-shot and many-shot fusion learning for mobile visual food recognition, it has a compact framework and is able to learn from existing dataset categories, and also new food categories given only a few sample images. We construct a new Indian food dataset called NTU-IndianFood107 in order to evaluate the performance of the proposed method. The dataset has two parts: (i) a Base Dataset of 83 classes of Indian food images with over 600 images per class to perform many-shot learning, and (ii) a Food Diary of 24 classes captured in restaurants with limited number to simulate the few-shot learning on new food categories. The proposed fusion method achieves a Top-1 classification accuracy of 72.0% on the new dataset. Kim-Hui Yap, Alex Chichung Kot, Ling-Yu Duan, Ngai-Man Cheung |
ISCAS | 4 |
| 2019 | Lossy Intermediate Deep Learning Feature Compression and EvaluationabstractWith the unprecedented success of deep learning in computer vision tasks, many cloud-based visual analysis applications are powered by deep learning models. However, the deep learning models are also characterized with high computational complexity and are task-specific, which may hinder the large-scale implementation of the conventional data communication paradigms. To enable a better balance among bandwidth usage, computational load and the generalization capability for cloud-end servers, we propose to compress and transmit intermediate deep learning features instead of visual signals and ultimately utilized features. The proposed strategy also provides a promising way for the standardization of deep feature coding. As the first attempt to this problem, we present a lossy compression framework and evaluation metrics for intermediate deep feature compression. Comprehensive experimental results show the effectiveness of our proposed methods and the feasibility of the proposed data transmission strategy. It is worth mentioning that the proposed compression framework and evaluation metrics have been adopted into the ongoing AVS (Audio Video Coding Standard Workgroup) - Visual Feature Coding Standard. Zhuo Chen 0006, Kui Fan, Shiqi Wang 0001, Ling-Yu Duan, Weisi Lin, Alex Chichung Kot |
ACM Multimedia | 4 |
| 2019 | Market2Dish: A Health-aware Food Recommendation SystemabstractIn order to help people develop healthy eating habits, we present a personalized health-aware food recommendation system, calledMarket2Dish. Market2Dish could recognize the ingredients in the micro-videos taken from the market, characterize the health conditions of users from their social media accounts, and ultimately recommend users with the personalized healthy foods. Specifically, we employ a word-class interaction based text classification model to learn the fine-grained similarity between sparse health features on the social media platforms and pre-defined health concepts, and then a category-aware hierarchical memory network based recommender is introduced to learn the user-recipe interactions for better food recommendations. Moreover, we demonstrate this system as an online app for real-time interactions with users. Wenjie Wang 0007, Meng Liu 0006, Liqiang Nie, Ling-Yu Duan, Changsheng Xu |
ACM Multimedia | 5 |
| 2019 | Adaptive Feature Fusion via Graph Neural Network for Person Re-identificationabstractPerson Re-identification (ReID) targets to identify a probe person appeared under multiple camera views. Existing methods focus on proposing a robust model to capture the discriminative information. However, they all generate a representation by mining useful clues from a given single image, and ignore the intercommunication with other images. To address this issue, we propose a novel network named Feature-Fusing Graph Neural Network (FFGNN), which fully utilizes the relationships among the nearest neighbors of the given image, and allows message propagation to update the feature of the node during representation learning. Given an anchor image, the FFGNN firstly obtains its Top-K nearest images based on the feature generated by the trained Feature-Extracting Network(FEN). We then construct a graph G based on the obtained K+1 images, in which each node represents the feature of an image. The edge of the graph G is obtained by combing the visual similarity and Jaccard similarity between nodes. Within the constructed graph G, FFGNN conducts message propagation and adaptive feature fusion between nodes by iteratively performing graph convolutional operation on the input features. Finally, the FFGNN outputs a robust and discriminative representation which contains the information from its similar images. Extensive experiments on three public person ReID datasets including Market-1501, DukeMTMC-ReID, and CUHK03 demonstrate that the proposed model can achieve significant improvement against state-of-the-art methods. Hantao Yao, Ling-Yu Duan, Hanxing Yao, Changsheng Xu |
ACM Multimedia | 3 |
| 2019 | See Through the Windshield from Surveillance CameraabstractThis paper attempts to address the challenging task of seeing through the windshield images captured by surveillance cameras in the wild. Such images usually have very low visibility due to heterogeneous degradations caused by blur, haze, reflection, noise etc., which makes existing image enhancing methods inapplicable. We propose a windshield image restoration generative adversarial network (WIRE-GAN) to restore and enhance the visibility of windshield images. We adopt the weakly supervised framework based on the generative model, which has effectively released the request of paired training data for a specific type of degradation. To generate more semantically consistent results even in extreme lighting conditions, we introduce a novel content-preserving strategy into the proposed weakly-supervised framework. To make the image restoration more reliable, the WIRE-GAN network constructs a sort of content-aware embedding space and enforces the constraint of the restored windshield images being closer to the original input in the embedding space. Moreover, we collect a large-scale windshield image dataset (WIRE dataset) to validate the advantage of our method in improving the image quality, and further evaluate the impact of windshield restoration on the vehicle ReID performance. Daiqian Ma, Renjie Wan, Ce Wang 0007, Boxin Shi, Ling-Yu Duan |
ACM Multimedia | 6 |
| 2019 | Toward Intelligent Visual Sensing and Low-cost Analysis: A Collaborative Computing ApproachabstractIn the big data era, there has been an increasing consensus that the label information, computational resources and communication bandwidth are particularly precious. State-of-the-art research is revolutionizing the vision systems of the smart city, which converts the visual signals from sensory input into feature representations and conveys the compact feature for analysis by using the computational resources of both front and back ends. To deploy a robust model, large amounts of labeled data are usually required, and thereby heavy computational and communication resources are incurred in model training as well as inference. However, the computational resources in front-end devices are usually constrained, and heavy transmission burden is imposed when leveraging multiple models amongst different ends. In this work, we propose a novel collaborative computing approach for intelligent sensing and low-cost analysis, which reduces the requirement of labeled data and communication cost, and balances the computational load in model training and inference. By incorporating the adversarial learning mechanism into collaborative model training, knowledge of different domains can be better exploited. Moreover, the learned models are deployed for inference in a collaborative manner, in which part of model is placed in front-ends for extracting intermediate feature maps, and part of the model remains in back ends for inference with received feature maps. The effectiveness of the proposed approach has been validated in the context of an emerging digital retina system for smart city intelligent applications. Ling-Yu Duan, Yong Luo 0002, Shiqi Wang 0001, Yonggang Wen 0001, Wen Gao 0001 |
VCIP | 2 |
| 2019 | DeepShoe: An improved Multi-Task View-invariant CNN for street-to-shop shoe retrieval
Huijing Zhan, Boxin Shi, Ling-Yu Duan, Alex Chichung Kot |
Comput. Vis. Image Underst. | 3 |
| 2019 | IDeRs: Iterative dehazing method for single remote sensing image
Long Xu 0001, Dong Zhao 0016, Yihua Yan, Sam Kwong, Jie Chen 0006, Ling-Yu Duan |
Inf. Sci. | 6 |
| 2019 | Toward Knowledge as a Service Over Networks: A Deep Learning Model Communication ParadigmabstractThe advent of artificial intelligence and Internet of Things has led to the seamless transition turning the big data into the big knowledge. The deep learning models, which assimilate knowledge from large-scale data, can be regarded as an alternative but promising modality of knowledge for artificial intelligence services. Yet, the compression, storage, and communication of the deep learning models towards better knowledge services, especially over networks, pose a set of challenging problems on both industrial and academic realms. This paper presents the deep learning model communication paradigm based on multiple model compression, which greatly exploits the redundancy among multiple deep learning models in different application scenarios. We analyze the potential and demonstrate the promise of the compression strategy for deep learning model communication through a set of experiments. Moreover, the interoperability in deep learning model communication, which is enabled based on the standardization of compact deep learning model representation, is also discussed and envisioned. Ziqian Chen, Ling-Yu Duan, Shiqi Wang 0001, Yihang Lou, Tiejun Huang 0001, Dapeng Oliver Wu, Wen Gao 0001 |
IEEE J. Sel. Areas Commun. | 2 |
| 2019 | Front-End Smart Visual Sensing and Back-End Intelligent Analysis: A Unified Infrastructure for Economizing the Visual System of City BrainabstractThe visual data, which are acquired from the ubiquitous visual sensors deployed in metropolitans, are of great value and paramount significance to enhance the effectiveness and pursue the future development of smart cities. In this paper, the essential building blocks of the unified visual data management and analysis infrastructure that serve as the foundation for the economical visual system in the city brain, are introduced to facilitate the utilization of the visual signal in the artificial intelligence era. In particular, we start by the discussion of the front-end smart visual sensing in the context of economical communication and service with the heterogeneous network, and the functionalities and necessities of compact visual feature and deep learning model representations are detailed. Subsequently, the utilities of the infrastructure are demonstrated through two intelligent applications at the back-end, including vehicle re-identification and person re-identification. The standardizations regarding compact feature and deep neural network representations, which are regarded as the key ingredients in this infrastructure and greatly facilitate the construction of the visual system in the city brain, are also discussed. Finally, we envision how the potential issues regarding the economical visual communications for future smart cities might be pragmatically approached within this unified infrastructure. Yihang Lou, Ling-Yu Duan, Shiqi Wang 0001, Ziqian Chen, Chang Wen Chen, Wen Gao 0001 |
IEEE J. Sel. Areas Commun. | 2 |
| 2019 | Learning to remove reflections from windshield images
Ce Wang 0007, Boxin Shi, Ling-Yu Duan |
Signal Process. Image Commun. | 3 |
| 2019 | Multi-scale Optimal Fusion model for single image dehazing
Dong Zhao 0016, Long Xu 0001, Yihua Yan, Jie Chen 0006, Ling-Yu Duan |
Signal Process. Image Commun. | 5 |
| 2019 | Robust Distracter-Resistive Tracker via Learning a Multi-Component Discriminative DictionaryabstractDiscriminative dictionary learning (DDL) provides an appealing paradigm for appearance modeling in visual tracking. However, most existing DDL-based trackers cannot handle drastic appearance changes, especially for scenarios with background cluster and/or similar object interference. One reason is that they often suffer from the loss of subtle visual information, which is critical to distinguish an object from distracters. In this paper, we explore the use of activations from the convolutional layer of a convolutional neural network to improve the object representation and then propose a robust distracter-resistive tracker via learning a multi-component discriminative dictionary. The proposed method exploits both the intra-class and inter-class visual information to learn shared atoms and the class-specific atoms. By imposing several constraints into the objective function, the learned dictionary is reconstructive, compressive, and discriminative, and thus can better distinguish an object from the background. In addition, our convolutional features have structural information for object localization and balance the discriminative power and semantic information of the object. Tracking is carried out within a Bayesian inference framework where a joint decision measure is used to construct the observation model. To alleviate the drift problem, the reliable tracking results obtained online are accumulated to update the dictionary. Both the qualitative and quantitative results on the CVPR2013 benchmark, the VOT2015 data set, and the SPOT data set demonstrate that our tracker achieves substantially better overall performance against the state-of-the-art approaches. Weichao Shen, Yuwei Wu 0001, Junsong Yuan 0001, Ling-Yu Duan, Jian Zhang 0002, Yunde Jia |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Embedding Adversarial Learning for Vehicle Re-IdentificationabstractThe high similarities of different real-world vehicles and great diversities of the acquisition views pose grand challenges to vehicle re-identification (ReID), which traditionally maps the vehicle images into a high-dimensional embedding space for distance optimization, vehicle discrimination, and identification. To improve the discriminative capability and robustness of the ReID algorithm, we propose a novel end-to-end embedding adversarial learning network (EALN) that is capable of generating samples localized in the embedding space. Instead of selecting abundant hard negatives from the training set, which is extremely difficult if not impossible, with our embedding adversarial learning scheme, the automatically generated hard negative samples in the specified embedding space can greatly improve the capability of the network for discriminating similar vehicles. Moreover, the more challenging cross-view vehicle ReID problem, which requires the ReID algorithm to be robust with different query views, can also benefit from such a scheme based on the artificially generated cross-view samples. We demonstrate the promise of EALN through extensive experiments and show the effectiveness of hard negative and cross-view generation in facilitating vehicle ReID based on the comparisons with the state-of-the-art schemes. Yihang Lou, Jun Liu 0036, Shiqi Wang 0001, Ling-Yu Duan |
IEEE Trans. Image Process. | 5 |
| 2019 | Unified Spatio-Temporal Attention Networks for Action Recognition in VideosabstractRecognizing actions in videos is not a trivial task because video is an information-intensive media and includes multiple modalities. Moreover, on each modality, an action may only appear at some spatial regions, or only part of the temporal video segments may contain the action. A valid question is how to locate the attended spatial areas and selective video segments for action recognition. In this paper, we devise a general attention neural cell, called AttCell, that estimates the attention probability not only at each spatial location but also for each video segment in a temporal sequence. With AttCell, a unified Spatio-Temporal Attention Networks (STAN) is proposed in the context of multiple modalities. Specifically, STAN extracts the feature map of one convolutional layer as the local descriptors on each modality and pools the extracted descriptors with the spatial attention measured by AttCell as a representation of each segment. Then, we concatenate the representation on each modality to seek a consensus on the temporal attention, a priori, to holistically fuse the combined representation of video segments to the video representation for recognition. Our model differs from conventional deep networks, which focus on the attention mechanism, because our temporal attention provides a principled and global guidance across different modalities and video segments. Extensive experiments are conducted on four public datasets; UCF101, CCV, THUMOS14, and Sports-1M; our STAN consistently achieves superior results over several state-of-the-art techniques. More remarkably, we validate and demonstrate the effectiveness of our proposal when capitalizing on the different number of modalities. Dong Li 0019, Ting Yao 0003, Ling-Yu Duan, Tao Mei 0001, Yong Rui |
IEEE Trans. Multim. | 3 |
| 2019 | Codebook-Free Compact Descriptor for Scalable Visual SearchabstractThe MPEG compact descriptors for visual search (CDVS) is a standard toward image matching and retrieval. To achieve high retrieval accuracy over a large scale image/video dataset, recent research efforts have demonstrated that employing extremely high-dimensional descriptors such as the Fisher vector (FV) and the vector of locally aggregated descriptors (VLAD) can yield good performance. Since the FV (or VLAD) possesses high discriminability but small visual vocabulary, it has been adopted by CDVS to construct a global compact descriptor. In this paper, we study the development of global compact descriptors in the completed CDVS standard and the emerging compact descriptors for video analysis (CDVA) standard, in which we formulate the FV (or VLAD) compression as a resource-constrained optimization problem. Accordingly, we propose a codebook-free aggregation method via dual selection to generate a global compact visual descriptor, which supports fast and accurate feature matching free of large visual codebooks, fulfilling the low memory requirement of mobile visual search at significantly reduced latency. Specifically, we investigate both sample-specific Gaussian component redundancy and bit dependency within a binary aggregated descriptor to produce compact binary codes. Our technique contributes to the scalable compressed Fisher vector (SCFV) adopted by the CDVS standard. Moreover, the SCFV descriptor is currently serving as the frame-level hand-crafted video feature, which inspires the inheritance of CDVS descriptors for the emerging CDVA standard. Furthermore, we investigate the positive complementary effect of our standard compliant compact descriptor and deep learning based features extracted from convolutional neural networks with significant mean average precision gains. Extensive evaluation over benchmark databases shows the significant merits of the codebook-free binary codes for scalable visual search. Yuwei Wu 0001, Feng Gao 0014, Jie Lin 0001, Vijay Chandrasekhar 0001, Junsong Yuan 0001, Ling-Yu Duan |
IEEE Trans. Multim. | 7 |
| 2018 | SSNet: Scale Selection Network for Online 3D Action PredictionabstractIn action prediction (early action recognition), the goal is to predict the class label of an ongoing action using its observed part so far. In this paper, we focus on online action prediction in streaming 3D skeleton sequences. A dilated convolutional network is introduced to model the motion dynamics in temporal dimension via a sliding window over the time axis. As there are significant temporal scale variations of the observed part of the ongoing action at different progress levels, we propose a novel window scale selection scheme to make our network focus on the performed part of the ongoing action and try to suppress the noise from the previous actions at each time step. Furthermore, an activation sharing scheme is proposed to deal with the overlapping computations among the adjacent steps, which allows our model to run more efficiently. The extensive experiments on two challenging datasets show the effectiveness of the proposed action prediction framework. Jun Liu 0036, Amir Shahroudy, Gang Wang 0012, Ling-Yu Duan, Alex Chichung Kot |
CVPR | 4 |
| 2018 | CRRN: Multi-Scale Guided Concurrent Reflection Removal NetworkabstractRemoving the undesired reflections from images taken through the glass is of broad application to various computer vision tasks. Non-learning based methods utilize different handcrafted priors such as the separable sparse gradients caused by different levels of blurs, which often fail due to their limited description capability to the properties of real-world reflections. In this paper, we propose the Concurrent Reflection Removal Network (CRRN) to tackle this problem in a unified framework. Our proposed network integrates image appearance information and multi-scale gradient information with human perception inspired loss function, and is trained on a new dataset with 3250 reflection images taken under diverse real-world scenes. Extensive experiments on a public benchmark dataset show that the proposed method performs favorably against state-of-the-art methods. Renjie Wan, Boxin Shi, Ling-Yu Duan, Ah-Hwee Tan, Alex Chichung Kot |
CVPR | 3 |
| 2018 | Gated Square-Root Pooling for Image Instance RetrievalabstractRecently Convolutional Neural Networks (CNNs) have achieved great success in different fields including image instance retrieval. However traditional global pooling approaches fail to capture all possible discriminative information of CNN activations and treat activations over channels equally regardless of the different importance between channels. In this work, we focus on the mentioned problem of global feature pooling over CNN activations for image instance retrieval. We make two contributions. First, we introduce a channel-wise SQUare-root (SQU) pooling (2-norm) approach, which makes better use of information over activation maps and is superior to Average (1-norm) and Max pooling (infinity norm), in the context of instance retrieval. Second, we further improve SQU by learning a gating function that weights the contributions of different channels, in an end-to-end manner. Extensive experiments on 6 benchmark datasets show that the proposed strategies achieve considerable improvements over state-of-the-art. Ziqian Chen, Jie Lin 0001, Vijay Chandrasekhar 0001, Ling-Yu Duan |
ICIP | 4 |
| 2018 | From Data to Knowledge: Deep Learning Model Compression, Transmission and CommunicationabstractWith the advances of artificial intelligence, recent years have witnessed a gradual transition from the big data to the big knowledge. Based on the knowledge-powered deep learning models, the big data such as the vast text, images and videos can be efficiently analyzed. As such, in addition to data, the communication of knowledge implied in the deep learning models is also strongly desired. As a specific example regarding the concept of knowledge creation and communication in the context of Knowledge Centric Networking (KCN), we investigate the deep learning model compression and demonstrate its promise use through a set of experiments. In particular, towards future KCN, we introduce efficient transmission of deep learning models in terms of both single model compression and multiple model prediction. The necessity, importance and open problems regarding the standardization of deep learning models, which enables the interoperability with the standardized compact model representation bitstream syntax, are also discussed. Ziqian Chen, Shiqi Wang 0001, Dapeng Oliver Wu, Tiejun Huang 0001, Ling-Yu Duan |
ACM Multimedia | 5 |
| 2018 | ChipGAN: A Generative Adversarial Network for Chinese Ink Wash Painting Style TransferabstractStyle transfer has been successfully applied on photos to generate realistic western paintings. However, because of the inherently different painting techniques adopted by Chinese and western paintings, directly applying existing methods cannot generate satisfactory results for Chinese ink wash painting style transfer. This paper proposes ChipGAN, an end-to-end Generative Adversarial Network based architecture for photo to Chinese ink wash painting style transfer. The core modules of ChipGAN enforce three constraints -- voids, brush strokes, and ink wash tone and diffusion -- to address three key techniques commonly adopted in Chinese ink wash painting. We conduct stylization perceptual study to score the similarity of generated paintings to real paintings by consulting with professional artists based on the newly built Chinese ink wash photo and image dataset. The advantages in visual quality compared with state-of-the-art networks and high stylization perceptual study scores show the effectiveness of the proposed method. Feng Gao 0014, Daiqian Ma, Boxin Shi, Ling-Yu Duan |
ACM Multimedia | 5 |
| 2018 | A Unified Generative Adversarial Framework for Image Generation and Person Re-identificationabstractPerson re-identification (re-id) aims to match a certain person across multiple non-overlapping cameras. It is a challenging task because the same person's appearance can be very different across camera views due to the presence of large pose variations. To overcome this issue, in this paper, we propose a novel unified person re-id framework by exploiting person poses and identities jointly for simultaneous person image synthesis under arbitrary poses and pose-invariant person re-identification. The framework is composed of a GAN based network and two Feature Extraction Networks (FEN), and enjoys following merits. First, it is a unified generative adversarial model for person image generation and person re-identification. Second, a pose estimator is utilized into the generator as a supervisor in the training process, which can effectively help pose transfer and guide the image generation with any desired pose. As a result, the proposed model can automatically generate a person image under an arbitrary pose. Third, the identity-sensitive representation is explicitly disentangled from pose variations through the person identity and pose embedding. Fourth, the learned re-id model can have better generalizability on a new person re-id dataset by using the synthesized images as auxiliary samples. Extensive experimental results on four standard benchmarks including Market-1501 [69], DukeMTMC-reID [40], CUHK03 [23], and CUHK01 [22] demonstrate that the proposed model can perform favorably against state-of-the-art methods. Tianzhu Zhang 0001, Ling-Yu Duan, Changsheng Xu |
ACM Multimedia | 3 |
| 2018 | Multi-Scale Context Attention Network for Image RetrievalabstractRecent attempts on the Convolutional Neural Network (CNN) based image retrieval usually adopt the output of a specific convolutional or fully connected layer as feature representation. Though superior representation capability has yielded better retrieval performance, the scale variation and clutter distracting remain to be two challenging problems in CNN based image retrieval. In this work, we propose a Multi-Scale Context Attention Network (MSCAN) to generate global descriptors, which is able to selectively focus on the informative regions with the assistance of multi-scale context information. We model the multi-scale context information by an improved Long Short-Term Memory (LSTM) network across different layers. As such, the proposed global descriptor is equipped with the scale aware attention capability. Experimental results show that our proposed method can effectively capture the informative regions in images and retain reliable attention responses when encountering scale variation and clutter distracting. Moreover, we compare the performance of the proposed scheme with the state-of-the-art global descriptors, and extensive results verify that the proposed MSCAN can achieve superior performance on several image retrieval benchmarks. Yihang Lou, Shiqi Wang 0001, Ling-Yu Duan |
ACM Multimedia | 4 |
| 2018 | Depth Structure Preserving Scene Image GenerationabstractKey to automatically generate natural scene images is to properly arrange amongst various spatial elements, especially in the depth cue. To this end, we introduce a novel depth structure preserving scene image generation network (DSP-GAN), which favors a hierarchical architecture, for the purpose of depth structure preserving scene image generation. The main trunk of the proposed infrastructure is built upon a Hawkes point process that models high-order spatial dependency between different depth layers. Within each layer generative adversarial sub-networks are trained collaboratively to generate realistic scene components, conditioned on the layer information produced by the point process. We experiment our model on annotated natural scene images collected from SUN dataset and demonstrate that our models are capable of generating depth-realistic natural scene image. Wendong Zhang 0002, Feng Gao 0014, Bingbing Ni, Ling-Yu Duan, Yichao Yan, Jingwei Xu 0005, Xiaokang Yang 0001 |
ACM Multimedia | 4 |
| 2018 | Facial Expression Recognition in the Wild: A Cycle-Consistent Adversarial Attention Transfer ApproachabstractFacial expression recognition (FER) is a very challenging problem due to different expressions under arbitrary poses. Most conventional approaches mainly perform FER under laboratory controlled environment. Different from existing methods, in this paper, we formulate the FER in the wild as a domain adaptation problem, and propose a novel auxiliary domain guided Cycle-consistent adversarial Attention Transfer model (CycleAT) for simultaneous facial image synthesis and facial expression recognition in the wild. The proposed model utilizes large-scale unlabeled web facial images as an auxiliary domain to reduce the gap between source domain and target domain based on generative adversarial networks (GAN) embedded with an effective attention transfer module, which enjoys several merits. First, the GAN-based method can automatically generate labeled facial images in the wild through harnessing information from labeled facial images in source domain and unlabeled web facial images in auxiliary domain. Second, the class-discriminative spatial attention maps from the classifier in source domain are leveraged to boost the performance of the classifier in target domain. Third, it can effectively preserve the structural consistency of local pixels and global attributes in the synthesized facial images through pixel cycle-consistency and discriminative loss. Quantitative and qualitative evaluations on two challenging in-the-wild datasets demonstrate that the proposed model performs favorably against state-of-the-art methods. Feifei Zhang 0001, Tianzhu Zhang 0001, Qirong Mao, Ling-Yu Duan, Changsheng Xu |
ACM Multimedia | 4 |
| 2018 | Tracklet Siamese Network with Constrained Clustering for Multiple Object TrackingabstractMultiple object tracking (MOT) is an important yet challenging task in video understanding and analysis. Basically, MOT aims to associate detected objects into trajectories based on their temporal relationships. The occlusion among moving objects poses a major challenge towards robust modeling of these relationships. In this paper, we propose a novel Tracklet Siamese Network (TSN) for learning similarities between track-lets characterized by appearance information, achieving superior performance on two MOTChallenge benchmark datasets. Our framework constructs short tracklets from highly-related object detections by excluding inaccurate object detections. We also adopt a constrained clustering technique to piece tracklets together into long trajectories, thus recovering many missing detections caused by original detector or the detection removing in the previous step. Comparisons against state-of-the-art methods were reported while ablation studies further substantiate the viability of components in our approach. Jinlong Peng, Fan Qiu, John See, Shaoshuai Huang, Ling-Yu Duan, Weiyao Lin |
VCIP | 6 |
| 2018 | Rate-Distortion Optimized Sparse Coding With Ordered Dictionary for Image Set CompressionabstractImage set compression has recently emerged as an active research topic due to the rapidly increasing demand in cloud storage. In this paper, we propose a novel framework for image set compression based on the rate-distortion optimized sparse coding. Specifically, given a set of similar images, one representative image is first identified according to the similarity among these images, and a dictionary can be learned subsequently in wavelet domain from the training samples collected from the representative image. In order to improve coding efficiency, the dictionary atoms are reordered according to their use frequencies when representing the representative image. As such, the remaining images can be efficiently compressed with sparse coding based on the reordered dictionary that is highly adaptive to the content of the image set. To further improve the efficiency of sparse coding, the number of dictionary atoms for image patches is further optimized in a rate-distortion sense. Experimental results show that the proposed method can significantly improve the image compression performance compared with JPEG, JPEG2000, and the state-of-the-art dictionary learning-based methods. Xinfeng Zhang 0001, Weisi Lin, Yabin Zhang 0002, Shiqi Wang 0001, Siwei Ma 0001, Ling-Yu Duan, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2018 | Fast MPEG-CDVS Encoder With GPU-CPU Hybrid ComputingabstractThe compact descriptors for visual search (CDVS) standard from ISO/IEC moving pictures experts group has succeeded in enabling the interoperability for efficient and effective image retrieval by standardizing the bitstream syntax of compact feature descriptors. However, the intensive computation of a CDVS encoder unfortunately hinders its widely deployment in industry for large-scale visual search. In this paper, we revisit the merits of low complexity design of CDVS core techniques and present a very fast CDVS encoder by leveraging the massive parallel execution resources of graphics processing unit (GPU). We elegantly shift the computation-intensive and parallel-friendly modules to the state-of-the-arts GPU platforms, in which the thread block allocation as well as the memory access mechanism are jointly optimized to eliminate performance loss. In addition, those operations with heavy data dependence are allocated to CPU for resolving the extra but non-necessary computation burden for GPU. Furthermore, we have demonstrated the proposed fast CDVS encoder can work well with those convolution neural network approaches which enables to leverage the advantages of GPU platforms harmoniously, and yield significant performance improvements. Comprehensive experimental results over benchmarks are evaluated, which has shown that the fast CDVS encoder using GPU-CPU hybrid computing is promising for scalable visual search. Ling-Yu Duan, Wei Sun 0029, Xinfeng Zhang 0001, Shiqi Wang 0001, Jie Chen 0006, Jianxiong Yin, Simon See, Tiejun Huang 0001, Alex Chichung Kot, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Minimizing Reconstruction Bias Hashing via Joint Projection Learning and QuantizationabstractHashing, a widely-studied solution to the approximate nearest neighbor (ANN) search, aims to map data points in the high-dimensional Euclidean space to the low-dimensional Hamming space while preserving the similarity between original points. As directly learning binary codes can be NP-hard due to discrete constraints, a two-stage scheme, namely "projection and quantization", has already become a standard paradigm for learning similarity-preserving hash codes. However, most existing hashing methods typically separate these two stages and thus fail to investigate complementary effects of both stages. In this paper, we systematically study the relationship between "projection and quantization", and propose a novel minimal reconstruction bias hashing (MRH) method to learn compact binary codes, in which the projection learning and quantization optimizing are jointly performed. By introducing a lower bound analysis, we design an effective ternary search algorithm to solve the corresponding optimization problem. Furthermore, we conduct some insightful discussions on the proposed MRH approach, including the theoretical proof, and computational complexity. Distinct from previous works, MRH can adaptively adjust the projection dimensionality to balance the information loss between projection and quantization. The proposed framework not only provides a unique perspective to view traditional hashing methods but also evokes some other researches, e.g., guiding the design of the loss functions in deep networks. Extensive experiment results have shown that the proposed MRH significantly outperforms a variety of state-of-the-art methods over eight widely used benchmarks. Ling-Yu Duan, Yuwei Wu 0001, Zhe Wang 0019, Junsong Yuan 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Skeleton-Based Human Action Recognition With Global Context-Aware Attention LSTM NetworksabstractHuman action recognition in 3D skeleton sequences has attracted a lot of research attention. Recently, long short-term memory (LSTM) networks have shown promising performance in this task due to their strengths in modeling the dependencies and dynamics in sequential data. As not all skeletal joints are informative for action recognition, and the irrelevant joints often bring noise which can degrade the performance, we need to pay more attention to the informative ones. However, the original LSTM network does not have explicit attention ability. In this paper, we propose a new class of LSTM network, global context-aware attention LSTM, for skeleton-based action recognition, which is capable of selectively focusing on the informative joints in each frame by using a global context memory cell. To further improve the attention capability, we also introduce a recurrent attention mechanism, with which the attention performance of our network can be enhanced progressively. Besides, a two-stream framework, which leverages coarse-grained attention and fine-grained attention, is also introduced. The proposed method achieves state-of-the-art performance on five challenging datasets for skeleton-based action recognition. Jun Liu 0036, Gang Wang 0012, Ling-Yu Duan, Kamila Abdiyeva, Alex Chichung Kot |
IEEE Trans. Image Process. | 3 |
| 2018 | Region-Aware Reflection Removal With Unified Content and Gradient PriorsabstractRemoving the undesired reflections in images taken through the glass is of broad application to various image processing and computer vision tasks. Existing single image based solutions heavily rely on scene priors such as separable sparse gradients caused by different levels of blur, and they are fragile when such priors are not observed. In this paper, we notice that strong reflections usually dominant a limited region in the whole image, and propose a Region-aware Reflection Removal (R3) approach by automatically detecting and heterogeneously processing regions with and without reflections. We integrate content and gradient priors to jointly achieve missing contents restoration as well as background and reflection separation in a unified optimization framework. Extensive validation using 50 sets of real data shows that the proposed method outperforms state-of-the-art on both quantitative metrics and visual qualities. Renjie Wan, Boxin Shi, Ling-Yu Duan, Ah-Hwee Tan, Wen Gao 0001, Alex Chichung Kot |
IEEE Trans. Image Process. | 3 |
| 2018 | Group-Sensitive Triplet Embedding for Vehicle ReidentificationabstractThe widespread use of surveillance cameras toward smart and safe cities poses the critical but challenging problem of vehicle reidentification (Re-ID). The state-of-the-art research work performed vehicle Re-ID relying on deep metric learning with a triplet network. However, most existing methods basically ignore the impact of intraclass variance-incorporated embedding on the performance of vehicle reidentification, in which robust fine-grained features for large-scale vehicle Re-ID have not been fully studied. In this paper, we propose a deep metric learning method, group-sensitive-triplet embedding (GS-TRE), to recognize and retrieve vehicles, in which intraclass variance is elegantly modeled by incorporating an intermediate representation “group” between samples and each individual vehicle in the triplet network learning. To capture the intraclass variance attributes of each individual vehicle, we utilize an online grouping method to partition samples within each vehicle ID into a few groups, and build up the triplet samples at multiple granularities across different vehicle IDs as well as different groups within the same vehicle ID to learn fine-grained features. In particular, we construct a large-scale vehicle database “PKU-Vehicle,” consisting of 10 million vehicle images captured by different surveillance cameras in several cities, to evaluate the vehicle Re-ID performance in real-world video surveillance applications. Extensive experiments over benchmark datasets VehicleID, VeRI, and CompCar have shown that the proposed GS-TRE significantly outperforms the state-of-the-art approaches for vehicle Re-ID. Yihang Lou, Feng Gao 0014, Shiqi Wang 0001, Yuwei Wu 0001, Ling-Yu Duan |
IEEE Trans. Multim. | 6 |
| 2018 | Query Adaptive Multiview Object Instance Search and Localization Using SketchesabstractSketch-based object search is a challenging problem mainly due to three difficulties: 1) how to match the primary sketch query with the colorful image; 2) how to locate the small object in a big image that is similar to the sketch query; and 3) given the large image database, how to ensure an efficient search scheme that is reasonably scalable. To address the above challenges, we propose leveraging object proposals for object search and localization. However, instead of purely relying on sketch features, we propose fully utilizing the appearance features of object proposals to resolve the ambiguities between the matching sketch query and object proposals. Our proposed query adaptive search is formulated as a subgraph selection problem, which can be solved by the maximum flow algorithm. By performing query expansion, it can accurately locate the small target objects in a cluttered background or densely drawn deformation-intensive cartoon (Manga like) images. To improve the computing efficiency of matching proposal candidates, the proposed Multi View Spatially Constrained Proposal Selection encodes each identified object proposal in terms of a small local basis of anchor objects. The results on benchmark datasets validate the advantages of utilizing both the sketch and appearance features for sketch-based search, while ensuring sufficient scalability at the same time. Sreyasee Das Bhattacharjee, Junsong Yuan 0001, Jingjing Meng, Ling-Yu Duan |
IEEE Trans. Multim. | 5 |
| 2018 | Toward Intelligent Product Retrieval for TV-to-Online (T2O) Application: A Transfer Metric Learning ApproachabstractIt is desired (especially for young people) to shop for the same or similar products shown in the multimedia contents (such as online TV programs). This indicates an urgent demand for improving the experience of TV-to-Online (T2O). In this paper, a transfer learning approach as well as a prototype system for effortless T2O experience is developed. In the system, a key component is high-precision product search, which is to fulfill exact matching between a query item and the database ones. The matching performance primarily relies on distance estimation, but the data characteristics cannot be well modeled and exploited by a simple Euclidean distance. This motivates us to introduce distance metric learning (DML) for improving the distance estimation. However, in traditional DML methods, the side information (such as the similar/dissimilar constraints or relevance/irrelevance judgements) in the target domain is leveraged. These methods may fail due to limited side information. Fortunately, this issue can be alleviated by utilizing transfer metric learning (TML) to exploit information from other related domains. In this paper, a novel manifold regularized heterogeneous multitask metric learning framework is proposed, in which each domain is treated equally. The proposed approach allows us to simultaneously exploit the information from other domains and the unlabeled information. Furthermore, the ranking-based loss is adopted to make our model more appropriate for search. Experiments on two challenging real-world datasets demonstrate the effectiveness of the proposed method. This TML approach is expected to impact the transformation of the emerging T2O trend in both TV and online video domains. Qiang Fu 0006, Yong Luo 0002, Yonggang Wen 0001, Dacheng Tao, Ying Li 0012, Ling-Yu Duan |
IEEE Trans. Multim. | 6 |
| 2018 | Data-Driven Lightweight Interest Point Selection for Large-Scale Visual SearchabstractWith the explosive increase of images and videos, visual analysis has become an essential technique in dealing with the big visual data, which utilizes the visual feature descriptors to search or recognize the images or frames with target objects or events. Subject to the constraints of resources (e.g., memory, bandwidth, storage, etc.), interest point selection is crucial to generate robust compact descriptors for high-efficiency visual analysis by selecting and aggregating the most discriminative local feature descriptors, which has been demonstrated in the state-of-the-art low bit rate visual search works. In this paper, we propose a data-driven lightweight interest point selection approach to significantly improve the performance of visual search, while ameliorating the efficiency of extracting feature descriptors. Comprehensive experimental results over benchmarks have shown that the proposed interest point selection algorithm has significantly improved image matching and retrieval performance in the completed MPEG Compact Descriptors for Visual Search (CDVS) standard as well as the emerging MPEG Compact Descriptors for Video Analytics (CDVA) standard, say 20% mAP gain by data-driven selection against random selection of interest points. In particular, the presented data-driven interest point selection has been adopted by MPEG-CDVS and MPEG-CDVA as a normative technique to improve the aggregation of handcrafted features, which has contributed to the combination of handcrafted features and deep learning (CNN) features as well. Feng Gao 0014, Xinfeng Zhang 0001, Yong Luo 0002, Xiaoming Li 0001, Ling-Yu Duan |
IEEE Trans. Multim. | 6 |
| 2017 | Global Context-Aware Attention LSTM Networks for 3D Action RecognitionabstractLong Short-Term Memory (LSTM) networks have shown superior performance in 3D human action recognition due to their power in modeling the dynamics and dependencies in sequential data. Since not all joints are informative for action analysis and the irrelevant joints often bring a lot of noise, we need to pay more attention to the informative ones. However, original LSTM does not have strong attention capability. Hence we propose a new class of LSTM network, Global Context-Aware Attention LSTM (GCA-LSTM), for 3D action recognition, which is able to selectively focus on the informative joints in the action sequence with the assistance of global contextual information. In order to achieve a reliable attention representation for the action sequence, we further propose a recurrent attention mechanism for our GCA-LSTM network, in which the attention performance is improved iteratively. Experiments show that our end-to-end network can reliably focus on the most informative joints in each frame of the skeleton sequence. Moreover, our network yields state-of-the-art performance on three challenging datasets for 3D action recognition. Jun Liu 0036, Gang Wang 0012, Ping Hu 0001, Ling-Yu Duan, Alex Chichung Kot |
CVPR | 4 |
| 2017 | Compression of Deep Neural Networks for Image Instance RetrievalabstractImage instance retrieval is the problem of retrieving images from a database which contain the same object. Convolutional Neural Network (CNN) based descriptors are becoming the dominant approach for generating global image descriptors for the instance retrieval problem. One major drawback of CNN-based global descriptors is that uncompressed deep neural network models require hundreds of megabytes of storage making them inconvenient to deploy in mobile applications or in custom hardware. In this work, we study the problem of neural network model compression focusing on the image instance retrieval task. We study quantization, coding, pruning and weight sharing techniques for reducing model size for the instance retrieval problem. We provide extensive experimental results on the trade-off between retrieval performance and model size for different types of networks on several data sets providing the most comprehensive study on this topic. We compress models to the order of a few MBs: two orders of magnitude smaller than the uncompressed models while achieving negligible loss in retrieval performance1. Vijay Chandrasekhar 0001, Jie Lin 0001, Qianli Liao, Olivier Morère, Antoine Veillard, Ling-Yu Duan, Tomaso A. Poggio |
DCC | 6 |
| 2017 | Compact Deep Invariant Descriptors for Video RetrievalabstractWith emerging demand for large-scale video analysis, the Motion Picture Experts Group (MPEG) initiated the Compact Descriptor for Video Analysis (CDVA) standardization in 2014. In this work, we develop novel deep-learning features and incorporate them into the well-established CDVA evaluation framework to study its effectiveness in video analysis. In particular, we propose a Nested Invariance Pooling (NIP) method to obtain compact and robust Convolutional Neural Network (CNNs) descriptors. The CNNs descriptors are generated by applying three different pooling operations to the feature maps of CNNs in a nested way towards rotation and scale invariant feature representation. In particular, the rational, advantages and performance on the combination of CNNs and handcrafted descriptors are provided to better investigate the complementary effects of deep learnt and handcrafted features. Extensive experimental results show that the proposed CNNs descriptors outperform both state-of-the-art CNNs descriptors and canonical handcrafted descriptors adopted in CDVA Experimental Model (CXM) with significant mAP gains of 11.3% and 4.7%, respectively. Moreover, the combination of NIP derived deep invariant descriptors and handcrafted descriptors not only fulfills the lowest bitrate budget of CDVA, but also significantly advances the performance of CDVA core techniques. Yihang Lou, Jie Lin 0001, Shiqi Wang 0001, Jie Chen 0006, Vijay Chandrasekhar 0001, Ling-Yu Duan, Tiejun Huang 0001, Alex Chichung Kot, Wen Gao 0001 |
DCC | 7 |
| 2017 | Benchmarking Single-Image Reflection Removal AlgorithmsabstractRemoving undesired reflections from a photo taken in front of a glass is of great importance for enhancing the efficiency of visual computing systems. Various approaches have been proposed and shown to be visually plausible on small datasets collected by their authors. A quantitative comparison of existing approaches using the same dataset has never been conducted due to the lack of suitable benchmark data with ground truth. This paper presents the first captured Single-image Reflection Removal dataset ‘SIR2’ with 40 controlled and 100 wild scenes, ground truth of background and reflection. For each controlled scene, we further provide ten sets of images under varying aperture settings and glass thicknesses. We perform quantitative and visual quality comparisons for four state-of-the-art single-image reflection removal algorithms using four error metrics. Open problems for improving reflection removal algorithms are discussed at the end. Renjie Wan, Boxin Shi, Ling-Yu Duan, Ah-Hwee Tan, Alex Chichung Kot |
ICCV | 3 |
| 2017 | Deep regional feature pooling for video matchingabstractIn this work, we study the problem of deep global descriptors for video matching with regional feature pooling. We aim to analyze the joint effect of ROI (Region of Interest) size and pooling moment on video matching performance. To this end, we propose to mathematically model the distribution of video matching function with a pooling function nested in. Matching performance can be estimated by the separability of these class-conditional distributions between matching and non-matching pairs. Empirical studies on the challenging MPEG CDVA dataset demonstrate that performance trends are consistent with the estimation and experimental results, though the theoretical model is largely simplified compared to video matching and retrieval in practice. Jie Lin 0001, Vijay Chandrasekhar 0001, Yihang Lou, Shiqi Wang 0001, Ling-Yu Duan, Tiejun Huang 0001, Alex Chichung Kot |
ICIP | 6 |
| 2017 | A Multi-Block N-ary trie structure for exact r-neighbour search in hamming spaceabstractThis paper proposes a novel algorithm to solve the exact r-neighbour search problem in Hamming space. Existing r-neighbour search methods typically adopt hash table to index binary codes. Given a query, existing approaches search the nearest neighbours by checking all buckets of a Hamming ball centered at the query. The problem is these methods spend most of search time visiting empty buckets (lookup misses). In this paper, we adopt trie structure to index binary codes. We consider several continuous bits of a binary string as a block and use it as an atomic indexing element in trie structure, which is efficient in access speed and memory usage. Our method searches the nearest neighbours of a query by utilizing the records of nodes in trie structure to avoid lookup misses. We name the proposed indexing structure as Multi-Block N-ary Trie (MBNT). A theoretical analysis is given to prove that MBNT has less time cost than other hash-based methods. Extensive results show that MBNT outperforms state-of-the-art algorithms on several large scale benchmarks. Ling-Yu Duan, Zhe Wang 0019, Jie Lin 0001, Vijay Chandrasekhar 0001, Tiejun Huang 0001 |
ICIP | 2 |
| 2017 | GPU Based fast MPEG-CDVS encoderabstractThe compact descriptors for visual search (CDVS) standard from ISO/IEC Moving Picture Experts Group (MPEG) has received increasing attentions due to its effectiveness in mobile visual search related applications. To explore the efficiency of CDVS in real-time applications, we implement the first optimized CDVS encoder based on the graphics processing unit (GPU). In particular, the CDVS feature extraction is implemented in a CPU and GPU collaborative architecture, and the most of the computation-intensive operations are transferred to GPU platform, which achieves significant speedup for CDVS compared with CPU implementation. Extensive evaluations on standard datasets provide promising results of the proposed scheme for real applications scenarios. Wei Sun 0029, Xinfeng Zhang 0001, Shiqi Wang 0001, Jie Chen 0006, Ling-Yu Duan |
ICIP | 5 |
| 2017 | Incorporating intra-class variance to fine-grained visual recognitionabstractFine-grained visual recognition aims to capture discriminative characteristics amongst visually similar categories. The state-of-the-art research work has significantly improved the fine-grained recognition performance by deep metric learning using triplet network. However, the impact of intra-category variance on the performance of recognition and robust feature representation has not been well studied. In this paper, we propose to leverage intra-class variance in metric learning of triplet network to improve the performance of fine-grained recognition. Through partitioning training images within each category into a few groups, we form the triplet samples across different categories as well as different groups, which is called Group Sensitive TRiplet Sampling (GS-TRS). Accordingly, the triplet loss function is strengthened by incorporating intra-class variance with GS-TRS, which may contribute to the optimization objective of triplet network. Extensive experiments over benchmark datasets CompCar and VehicleID show that the proposed GS-TRS has significantly outperformed state-of-the-art approaches in both classification and retrieval tasks. Yan Em, Feng Gao 0014, Yihang Lou, Shiqi Wang 0001, Tiejun Huang 0001, Ling-Yu Duan |
ICME | 6 |
| 2017 | Improving object detection with region similarity learningabstractObject detection aims to identify instances of semantic objects of a certain class in images or videos. The success of state-of-the-art approaches is attributed to the significant progress of object proposal and convolutional neural networks (CNNs). Most promising detectors involve multi-task learning with an optimization objective of softmax loss and regression loss. The first is for multi-class categorization, while the latter is for improving localization accuracy. However, few of them attempt to further investigate the hardness of distinguishing different sorts of distracting background regions (i.e., negatives) from true object regions (i.e., positives). To improve the performance of classifying positive object regions vs. a variety of negative background regions, we propose to incorporate triplet embedding into learning objective. The triplet units are formed by assigning each negative region to a meaningful object class and establishing class-specific negatives, followed by triplets construction. Over the benchmark PASCAL VOC 2007, the proposed triplet embedding has improved the performance of well-known Fas-tRCNN model with a mAP gain of 2.1%. In particular, the state-of-the-art approach OHEM can benefit from the triplet embedding and has achieved a mAP improvement of 1.2%. Feng Gao 0014, Yihang Lou, Shiqi Wang 0001, Tiejun Huang 0001, Ling-Yu Duan |
ICME | 6 |
| 2017 | DeepHash for Image Instance Retrieval: Getting Regularization, Depth and Fine-Tuning RightabstractThis work focuses on representing very high-dimensional global image descriptors using very compact 64-1024 bit binary hashes for instance retrieval. We propose DeepHash: a hashing scheme based on deep networks. Key to making DeepHash work at extremely low bitrates are three important considerations -- regularization, depth and fine-tuning -- each requiring solutions specific to the hashing problem. In-depth evaluation shows that our scheme outperforms state-of-the-art methods over several benchmark datasets for both Fisher Vectors and Deep Convolutional Neural Network features, by up to 8.5% over other schemes. The retrieval performance with 256-bit hashes is close to that of the uncompressed floating point features -- a remarkable 512x compression. Jie Lin 0001, Olivier Morère, Antoine Veillard, Ling-Yu Duan, Hanlin Goh, Vijay Chandrasekhar 0001 |
ICMR | 4 |
| 2017 | Nested Invariance Pooling and RBM Hashing for Image Instance RetrievalabstractThe goal of this work is the computation of very compact binary hashes for image instance retrieval. Our approach has two novel contributions. The first one is Nested Invariance Pooling (NIP), a method inspired from i-theory, a mathematical theory for computing group invariant transformations with feed-forward neural networks. NIP is able to produce compact and well-performing descriptors with visual representations extracted from convolutional neural networks. We specifically incorporate scale, translation and rotation invariances but the scheme can be extended to any arbitrary sets of transformations. We also show that using moments of increasing order throughout nesting is important. The NIP descriptors are then hashed to the target code size (32-256 bits) with a Restricted Boltzmann Machine with a novel batch-level regularization scheme specifically designed for the purpose of hashing (RBMH). A thorough empirical evaluation with state-of-the-art shows that the results obtained both with the NIP descriptors and the NIP+RBMH hashes are consistently outstanding across a wide range of datasets. Olivier Morère, Jie Lin 0001, Antoine Veillard, Ling-Yu Duan, Vijay Chandrasekhar 0001, Tomaso A. Poggio |
ICMR | 4 |
| 2017 | From Part to Whole: Who is Behind the Painting?abstractCompared with normal modalities, the representations of paintings are much more complex due to its large intra-class and small inter-class variation. This poses more difficulties in the task of authorship identification. In this paper, we propose a multi-task multi-range (MTMR) representation framework and try to resolve this issue in two ways. First, we investigate how to improve the representation through multi-task learning. Specifically, we attempt to optimize authorship identification with subtly correlated identification tasks such as style, genre and date. Second, in order to make the representation more comprehensive and reduce the information loss from image scaling, we propose a multi-range structure which is composed of local, regional and global representations. Experiments on the two most representative large-scale painting datasets, Rijksmuseum Challenge and Wikiart, have shown that our method significantly outperforms the existing methods. To give better understanding and provide more effective predictions, we utilize random forest as the feature ranking method to analyze the importance of different features and apply external knowledge matching to further examine the predictions. Moreover, the framework's effects of identifying the authorship are visualized on the paintings' artist-characteristic regions and t-SNE is further applied to perform artist-based cluster analysis. Extensive validation has demonstrated that the proposed framework yields superior performance in the chanllenging task of painting authorship identification. Daiqian Ma, Feng Gao 0014, Yihang Lou, Shiqi Wang 0001, Tiejun Huang 0001, Ling-Yu Duan |
ACM Multimedia | 7 |
| 2017 | HNIP: Compact Deep Invariant Representations for Video Matching, Localization, and RetrievalabstractWith emerging demand for large-scale video analysis, MPEG initiated the compact descriptor for video analysis (CDVA) standardization in 2014. Beyond handcrafted descriptors adopted by the current MPEG-CDVA reference model, we study the problem of deep learned global descriptors for video matching, localization, and retrieval. First, inspired by a recent invariance theory, we propose a nested invariance pooling (NIP) method to derive compact deep global descriptors from convolutional neural networks (CNNs), by progressively encoding translation, scale, and rotation invariances into the pooled descriptors. Second, our empirical studies have shown that a sequence of well designed pooling moments (e.g., max or average) may drastically impact video matching performance, which motivates us to design hybrid pooling operations via NIP (HNIP). HNIP has further improved the discriminability of deep global descriptors. Third, the technical merits and performance improvements by combining deep and handcrafted descriptors are provided to better investigate the complementary effects. We evaluate the effectiveness of HNIP within the well-established MPEG-CDVA evaluation framework. The extensive experiments have demonstrated that HNIP outperforms the state-of-the-art deep and canonical handcrafted descriptors with significant mAP gains of 5.5% and 4.7%, respectively. In particular the combination of HNIP incorporated and handcrafted global descriptors has significantly boosted the performance of CDVA core techniques with comparable descriptor size. Jie Lin 0001, Ling-Yu Duan, Shiqi Wang 0001, Yihang Lou, Vijay Chandrasekhar 0001, Tiejun Huang 0001, Alex Chichung Kot, Wen Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2016 | Affinity Preserving Quantization for Hashing: A Vector Quantization Approach to Compact Learn Binary CodesabstractHashing techniques are powerful for approximate nearest neighbour (ANN) search.Existing quantization methods in hashing are all focused on scalar quantization (SQ) which is inferior in utilizing the inherent data distribution.In this paper, we propose a novel vector quantization (VQ) method named affinity preserving quantization (APQ) to improve the quantization quality of projection values, which has significantly boosted the performance of state-of-the-art hashing techniques.In particular, our method incorporates the neighbourhood structure in the pre- and post-projection data space into vector quantization.APQ minimizes the quantization errors of projection values as well as the loss of affinity property of original space.An effective algorithm has been proposed to solve the joint optimization problem in APQ, and the extension to larger binary codes has been resolved by applying product quantization to APQ.Extensive experiments have shown that APQ consistently outperforms the state-of-the-art quantization methods, and has significantly improved the performance of various hashing techniques. Zhe Wang 0019, Ling-Yu Duan, Tiejun Huang 0001, Wen Gao 0001 |
AAAI | 2 |
| 2016 | Depth-based local feature selection for mobile visual searchabstractSelecting local features is crucial in generating robust compact descriptors for mobile visual search. The state-of-the-art MPEG Compact Descriptors for Visual Search (CDVS) standard has utilized the intrinsic characteristics (e.g., scale, orientation, peak, center distance, etc.) of interest points to select salient local features for selective aggregation and compression of local feature descriptors at different bit rates. In particular, the statistics of center distance was considered as an important attribute to select features in mobile visual search, which heavily relies on the assumption of a centralized object in a 2-dimensional query image. However, the ad-hoc assumption would probably fail to delineate query objects in a cluttered scene. In this paper, we propose to incorporate the depth cue to select local features. As most mobile phones are not yet equipped with depth sensor, we recover the disparity of local features through an auxiliary image to fast estimate the depth of a query image. The experiments have shown that, the incorporation of depth cue into feature selection can significantly improve the retrieval performance of the state-of-the-art CDVS compact descriptors at lower bit rates. For example, the mAP is improved from 84.5% to 88.6% at 512 bytes. Zhaoliang Liu, Ling-Yu Duan, Jie Chen 0006, Tiejun Huang 0001 |
ICIP | 2 |
| 2016 | Two-stage pooling of deep convolutional features for image retrievalabstractConvolutional Neural Network (CNN) based image representations have achieved high performance in image retrieval tasks. However, traditional CNN based global representations either provide high-dimensional features, which incurs large memory consumption and computing cost, or inadequately capture discriminative information in images, which degenerates the functionality of CNN features. To address those issues, we propose a two-stage partial mean pooling (PMP) approach to construct compact and discriminative global feature representations. The proposed PMP is meant to tackle the limits of traditional max pooling and mean (or average) pooling. By injecting the PMP pooling strategy into the CNN based patch-level mid-level feature extraction and representation, we have significantly improved the state-of-the-art retrieval performance over several common benchmark datasets. Tiancheng Zhi, Ling-Yu Duan, Tiejun Huang 0001 |
ICIP | 2 |
| 2016 | Smart query expansion scheme for CDVS based on illumination and key featuresabstractGiven a query image, retrieving images depicting the same object in a large scale database is becoming an urgent and challenging task. Recently, Compact Description for Visual Search (CDVS) is drafted by the ISO/IEC Moving Pictures Experts Group (MPEG) to support image retrieval applications, and it has been published as an international standard. Unfortunately, with regard to applications with hugely mutative illumination, perspective and noisy background, CDVS suffers from an inevitable performance loss. In this paper, firstly we introduce the query expansion to address performance loss caused by the scene complexity in CDVS. Secondly, a query expansion instance selection method based on illumination is proposed, which achieves better performance. Thirdly, we adopt a key feature matching score based weighted strategy in basic query expansion to improve retrieval performance. We evaluate our proposed methods on the Oxford (5K images) dataset and a reality traffic vehicle dataset (12K images), and the result shows that the proposed methods boost mean average precision (MAP) by 7% ∼ 10% in Oxford dataset and 7% ∼17% in vehicle dataset. Chuang Zhu, Huizhu Jia, Ling-Yu Duan, Jiawen Song, Wen Gao 0001 |
ICPR | 4 |
| 2016 | To Project More or to Quantize More: Minimize Reconstruction Bias for Learning Compact Binary Codes
Zhe Wang 0019, Ling-Yu Duan, Junsong Yuan 0001, Tiejun Huang 0001, Wen Gao 0001 |
IJCAI | 2 |
| 2016 | A Compact Binary Aggregated Descriptor via Dual Selection for Visual SearchabstractTo achieve high retrieval accuracy over a large scale image/video dataset, recent research efforts have demonstrated that employing extremely high-dimensional descriptors such as the Fisher Vector (FV) and the Vector of Locally Aggregated Descriptors (VLAD) can yield good performance. To enable fast search, the FV (or VLAD) is usually compressed by product quantization (PQ) or hashing. However, compressing high-dimensional descriptors via PQ or hashing may become intractable and infeasible due to both the storage and computation requirements for the linear/nonlinear projection of PQ or hashing methods. We develop a novel compact aggregated descriptor via dual selection for visual search. We utilize both sample-specific Gaussian component redundancy and bit dependency within a binary aggregated descriptor to produce its compact binary codes. The proposed method can effectively reduce the codesize of the raw aggregated descriptors, without degrading the search accuracy or introducing additional memory footprint. We demonstrate the significant advantages of the proposed binary codes in solving the approximate nearest neighbor (ANN) visual search problem. Experimental results on extensive datasets show that our method outperforms the state-of-the-art methods. Yuwei Wu 0001, Zhe Wang 0019, Junsong Yuan 0001, Ling-Yu Duan |
ACM Multimedia | 4 |
| 2016 | Overview of the MPEG-CDVS StandardabstractCompact descriptors for visual search (CDVS) is a recently completed standard from the ISO/IEC moving pictures experts group (MPEG). The primary goal of this standard is to provide a standardized bitstream syntax to enable interoperability in the context of image retrieval applications. Over the course of the standardization process, remarkable improvements were achieved in reducing the size of image feature data and in reducing the computation and memory footprint in the feature extraction process. This paper provides an overview of the technical features of the MPEG-CDVS standard and summarizes its evolution. Ling-Yu Duan, Vijay Chandrasekhar 0001, Jie Chen 0006, Jie Lin 0001, Zhe Wang 0019, Tiejun Huang 0001, Bernd Girod, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Query-Adaptive Small Object Search Using Object Proposals and Shape-Aware DescriptorsabstractWhile there has been a significant amount of work on object search and image retrieval, the focus has primarily been on establishing effective models for the whole images, scenes, and objects occupying a large portion of an image. In this paper, we propose to leverage object proposals to identify small and smooth-structured objects in a large image database. Unlike popular methods exploring a coarse image-level pairwise similarity, the search is designed to exploit the similarity measures at the proposal level. An effective graph-based query expansion strategy is designed to assess each of these better matched proposals against all its neighbors within the same image for a precise localization. Combined with a shape-aware feature descriptor EdgeBoW, a set of more insightful edge-weights and node-utility measures, the proposed search strategy can handle varying view angles, illumination conditions, deformation, and occlusion efficiently. Experiments performed on a number of other benchmark datasets show the powerful and superior generalization ability of this single integrated framework in dealing with both clutter-intensive real-life images and poor-quality binary document images at equal dexterity. Sreyasee Das Bhattacharjee, Junsong Yuan 0001, Yap-Peng Tan, Ling-Yu Duan |
IEEE Trans. Multim. | 4 |
| 2015 | Overview of the MPEG CDVS StandardabstractTowards mobile visual search, compact visual descriptors have been well advocated in both academic and industry endeavors. Moving Picture Experts Group (MPEG) initiated the remarkable Compact Descriptors for Visual Search (CDVS) standard activity in Jan. 2010 to push forward the frontiers of compact descriptors in mobile internet industry. In Oct. 2014, MPEG CDVS successfully entered the Final Draft of International Standard. CDVS made a series of significant breakthroughs in high performance and low complexity compact descriptors. In this paper, we give an overview of the MPEG CDVS standard, with emphasis on the development of the core techniques and their technical merits. Ling-Yu Duan, Tiejun Huang 0001, Wen Gao 0001 |
DCC | 1 |
| 2015 | Optimizing Binary Fisher Codes for Visual SearchabstractFisher vectors (FV) aggregated from local invariant features (e.g., SIFT) is one of the state-of-the-art descriptors for visual search, due to high discriminability but small visual vocabulary. Nevertheless, a high-dimensional FV needs to be compressed into a compact descriptor for light storage and high matching eficiency. In this paper, we formulate the FV compression as a resource-constrained optimization problem. Our goal is to maximize search performance subject to the constraints of descriptor compactness, compression complexity in terms of memory usage and time cost. Accordingly, we present a selective binary Fisher codes (SBFC) to compress the raw FV. Firstly, to fulfill the constraint of compression complexity, we binarize the FV by a sign function, Secondly, we propose to select discriminative bits from the binarized FV (BFC) to maximize search performance, subject to the constraint of descriptor compactness. Extensive experiments over MPEG Compact Descriptor for Visual Search (CDVS) benchmark datasets have shown that S-BFC significantly improves search performance at a smaller descriptor size as well as much lower complexity, compared with the state-of-the-art FV compression algorithms like Hashing and Product Quantziation (PQ). A simplified version of SBFC, SBFC LS has been adopted by the MPEG CDVS standard. In the CDVS evaluation framework, SBFC LS has achieved promising performance mean Average Precision (mAP) 83% on average at much lower memory cost of 40KB. Zhe Wang 0019, Ling-Yu Duan, Jie Lin 0001, Jie Chen 0006, Tiejun Huang 0001, Wen Gao 0001 |
DCC | 2 |
| 2015 | An efficient coding framework for compact descriptors extracted from video sequenceabstractTowards effective and efficient image matching or retrieval tasks, the emerging MPEG standard, named Compact Descriptors for Visual Search (CDVS), has fulfilled compact descriptors for still images, consisting of compressed local and global descriptor. Nevertheless, the frame-level coding of CDVS descriptors from a video sequence does not address the inter-frame redundancy issue, which may consume considerable bandwidth and storage resources. In this work, we propose an efficient coding framework of CDVS descriptors to generate compact descriptors for video sequences. For local descriptors, we propose a multiple reference predictive technique to exploit the temporal correlation of local descriptors and location coordinates over a sequence of frames. To further improve the prediction performance, keypoint tracking is applied to identify temporally repeated keypoints. For global descriptors, a propagation coding way is employed to compress the global descriptors of adjacent frames. The empirical evaluation has shown that the proposed coding approach has yielded a low bit rate of less than 40kbps on average, while maintaining comparable matching and retrieval performance. Compared to the sequence of original frame-level CDVS descriptors, the proposed approach has achieved over 25× bit rate reduction. Zhangshuai Huang, Ling-Yu Duan, Jie Lin 0001, Shiqi Wang 0001, Siwei Ma 0001, Tiejun Huang 0001 |
ICIP | 2 |
| 2015 | Hierarchical multi-VLAD for image retrievalabstractConstructing discriminative feature descriptors is crucial towards effective image retrieval. The state-of-the-art powerful global descriptor for this purpose is Vector of Locally Aggregated Descriptors (VLAD). Given a set of local features (say, SIFT) extracted from an image, the VLAD is generated by quantizing local features with a small visual vocabulary (64 to 512 centroids), aggregating the residual statistics of quantized features for each centroid and concatenating the aggregated residual vectors from each centroid. One can increase the search accuracy by increasing the size of vocabulary (from hundreds to hundreds of thousands), which, however, it leads to heavy computation cost with flat quantization. In this paper, we propose a hierarchical multi-VLAD to seek the tradeoff between descriptor discriminability and computation complexity. We build up a tree-structured hierarchical quantization (TSHQ) to accelerate the VLAD computation with a large vocabulary. As quantization error may propagate from root to leaf node (centroid) with TSHQ, we introduce multi-VLAD, which constructing a VLAD descriptor for each level of the vocabulary tree, so as to compensate for the quantization error at that level. Extensive evaluation over benchmark datasets has shown that the proposed approach outperforms state-of-the-art in terms of retrieval accuracy, fast extraction, as well as light memory cost. Ling-Yu Duan, Jie Lin 0001, Zhe Wang 0019, Tiejun Huang 0001 |
ICIP | 2 |
| 2015 | Hamming Compatible Quantization for Hashing
Zhe Wang 0019, Ling-Yu Duan, Jie Lin 0001, Tiejun Huang 0001, Wen Gao 0001 |
IJCAI | 2 |
| 2015 | Query-Adaptive Logo Search using Shape-Aware DescriptorsabstractWe propose a graph-based optimization framework to leverage category independent object proposals (candidate object regions) for logo search in a large scale image database. The proposed contour-based feature descriptor EdgeBoW is robust to view-angle changes, varying illumination conditions and can implicitly capture the significant object shape information. Having been equipped with a local descriptor, it can handle a fair amount of occlusion and deformation frequently present in a real-life scenario. Given a small set of initially retrieved candidate object proposals, a fast graph-based short-listing scheme is designed to exploit the mutual similarities among these proposals for eliminating outliers. In contrast to a coarse image-level pairwise similarity measure, this search focussed on a few specific image regions provides a more accurate method for matching. The proposed query expansion strategy aims to assess each of the remaining better matched proposals against all its neighbors within the same image for a precise localization. Combined with an efficient feature descriptor EdgeBoW, a set of more insightful edge-weights and node-utility measures can yield promising results, specially for object categories primarily defined by its shape. Extensive set of experiments performed on a number of benchmark datasets demonstrates its effectiveness and superior generalization ability in both clutter intensive real-life images and poor quality binary document images. Sreyasee Das Bhattacharjee, Junsong Yuan 0001, Yap-Peng Tan, Ling-Yu Duan |
ACM Multimedia | 4 |
| 2015 | Efficient image retrieval based mobile indoor localizationabstractVision based localization has been investigated for many years. The existing Structure from Motion (SfM) technique can reconstruct the 3D models based on the input images. The image retrieval and feature matching allow us to find the correspondence between the query image and the 3D model. According to these, the location can be easily calculated. In mobile scenarios, the limited CPU speed, memory storage and network latency bring in new challenges. The state-of-the-art solution can not be easily adopted due to the complicated calculation and large resource consumption. In this paper, we leverage the techniques developed during the MPEG-7 Compact Descriptors for Visual Search (CDVS) standardization, which aims to provide high performance and low complexity compact descriptors. We show that these techniques are suitable for mobile device and can achieve state-of-the-art retrieval performance in indoor environment. Besides, we propose additional components including blur measurement and result smoothing to improve the performance of the location calculation process. Based on these techniques, a whole system which enables fast vision based localization on mobile device is developed. We present experiments on the real world situation, showing that the system can strike a balance between accuracy and efficiency. Ruoyun He, Qingyi Tao, Jianfei Cai 0001, Ling-Yu Duan |
VCIP | 5 |
| 2015 | Finding the Secret of Image Saliency in the Frequency DomainabstractThere are two sides to every story of visual saliency modeling in the frequency domain. On the one hand, image saliency can be effectively estimated by applying simple operations to the frequency spectrum. On the other hand, it is still unclear which part of the frequency spectrum contributes the most to popping-out targets and suppressing distractors. Toward this end, this paper tentatively explores the secret of image saliency in the frequency domain. From the results obtained in several qualitative and quantitative experiments, we find that the secret of visual saliency may mainly hide in the phases of intermediate frequencies. To explain this finding, we reinterpret the concept of discrete Fourier transform from the perspective of template-based contrast computation and thus develop several principles for designing the saliency detector in the frequency domain. Following these principles, we propose a novel approach to design the saliency detector under the assistance of prior knowledge obtained through both unsupervised and supervised learning processes. Experimental results on a public image benchmark show that the learned saliency detector outperforms 18 state-of-the-art approaches in predicting human fixations. Jia Li 0003, Ling-Yu Duan, Xiaowu Chen 0001, Tiejun Huang 0001, Yonghong Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | A Low Complexity Interest Point DetectorabstractInterest point detection is a fundamental approach to feature extraction in computer vision tasks. To handle the scale invariance, interest points usually work on the scale-space representation of an image. In this letter, we propose a novel block-wise scale-space representation to significantly reduce the computational complexity of an interest point detector. Laplacian of Gaussian (LoG) filtering is applied to implement the block-wise scale-space representation. Extensive comparison experiments have shown the block-wise scale-space representation enables the efficient and effective implementation of an interest point detector in terms of memory and time complexity reduction, as well as promising performance in visual search. Jie Chen 0006, Ling-Yu Duan, Feng Gao 0014, Jianfei Cai 0001, Alex Chichung Kot, Tiejun Huang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2015 | Depth-Preserving Warping for Stereo Image RetargetingabstractThe popularity of stereo images and various display devices poses the need of stereo image retargeting techniques. Existing warping-based retargeting methods can well preserve the shape of salient objects in a retargeted stereo image pair. Nevertheless, these methods often incur depth distortion, since they attempt to preserve depth by maintaining the disparity of a set of sparse correspondences, rather than directly controlling the warping. In this paper, by considering how to directly control the warping functions, we propose a warping-based stereo image retargeting approach that can simultaneously preserve the shape of salient objects and the depth of 3D scenes. We first characterize the depth distortion in terms of warping functions to investigate the impact of a warping function on depth distortion. Based on the depth distortion model, we then exploit binocular visual characteristics of stereo images to derive region-based depth-preserving constraints which directly control the warping functions so as to faithfully preserve the depth of 3D scenes. Third, with the region-based depth-preserving constraints, we present a novel warping-based stereo image retargeting framework. Since the depth-preserving constraints are derived regardless of shape preservation, we relax the depth-preserving constraints to fulfill a tradeoff between shape preservation and depth preservation. Finally, we propose a quad-based implementation of the proposed framework. The results demonstrate the efficacy of our method in both depth and shape preservation for stereo image retargeting. Bing Li 0024, Ling-Yu Duan, Chia-Wen Lin, Tiejun Huang 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2015 | Weighted Component Hashing of Binary Aggregated Descriptors for Fast Visual SearchabstractTowards low bit rate mobile visual search, recent works have proposed to aggregate the local features and compress the aggregated descriptor (such as Fisher vector, the vector of locally aggregated descriptors) for low latency query delivery as well as moderate search complexity. Even though Hamming distance can be computed very fast, the computational cost of exhaustive linear search over the binary descriptors grows linearly with either the length of a binary descriptor or the number of database images. In this paper, we propose a novel weighted component hashing (WeCoHash) algorithm for long binary aggregated descriptors to significantly improve search efficiency over a large scale image database. Accordingly, the proposed WeCoHash has attempted to address two essential issues in Hashing algorithms: “what to hash” and “how to search.” “What to hash” is tackled by a hybrid approach, which utilizes both image-specific component (i.e., visual word) redundancy and bit dependency within each component of a binary aggregated descriptor to produce discriminative hash values for bucketing. “How to search” is tackled by an adaptive relevance weighting based on the statistics of hash values. Extensive comparison results have shown that WeCoHash is at least 20 times faster than linear search and 10 times faster than local sensitive hash (LSH) when maintaining comparable search accuracy. In particular , the WeCoHash solution has been adopted by the emerging MPEG compact descriptor for visual search (CDVS) standard to significantly speed up the exhaustive search of the binary aggregated descriptors. Ling-Yu Duan, Jie Lin 0001, Zhe Wang 0019, Tiejun Huang 0001, Wen Gao 0001 |
IEEE Trans. Multim. | 1 |
| 2014 | Region-based depth-preserving stereoscopic image retargetingabstractThe popularity of stereo images and various sizes of display screens pose the need of stereo image retargeting techniques which resize stereo image pairs to desired sizes. Many content-aware stereo image retargeting methods adapt the images through non-uniformly resizing regions. However, these methods often make the depth of retargeted version inconsistent with the original one, since they do not explicitly consider different effects of resizing distinct regions on the depths of 3D scenes. In this paper, we analyze the effects of region-wise resizing on the depths of 3D scenes. With such insights, we can properly edit or maintain the depth of a stereo image pair via region-wise resizing. In addition, by taking into account the effects on different regions, we propose a grid-based retargeting model for stereo images, which simultaneously preserve the depths of 3D scenes and the shapes of salient objects. Experimental results demonstrate the superior performance of our method. Bing Li 0024, Ling-Yu Duan, Chia-Wen Lin, Wen Gao 0001 |
ICIP | 2 |
| 2014 | Joint optimization of JPEG quantization table and coefficient thresholding for low bitrate mobile visual searchabstractLow latency query delivery over wireless network is a key problem for mobile visual search. Extracting compact descriptors directly on the mobile device is computational expensive, an alternate approach is to send highly compressed JPEG query images. As JPEG baseline optimizes the rate-distortion from a perceptual perspective rather than maintaining search performance, recent work proposed to learn a feature-preserving JPEG quantization table for improved search accuracy. However, this method is data-dependent and the quantization table cannot adapt to image blocks. To address these issues, we propose to jointly optimize the JPEG quantization table and coefficient thresholding. The matching score between uncompressed image and its compressed JPEG image is employed as the distortion measure to avoid time consuming image labeling, and coefficient thresholding eliminates the redundant coefficients. Extensive experiments on benchmark datasets show that our approach obtains superior performance than state-of-the-art at low bitrates, meanwhile, it consumes lower cost including processing time, memory and battery on mobile device. Ling-Yu Duan, Jie Lin 0001, Tiejun Huang 0001, Wen Gao 0001 |
ICIP | 2 |
| 2014 | Component hashing of variable-length binary aggregated descriptors for fast image searchabstractCompact locally aggregated binary features have shown great advantages in image search. As the exhaustive linear search in Hamming space still entails too much computational complexity for large datasets, recent works proposed to directly use binary codes as hash indices, yielding a dramatic increase in speedup. However, these methods cannot be directly applied to variable-length binary features. In this paper, we propose a Component Hashing (CoHash) algorithm to handle the variable-length binary aggregated descriptors indexing for fast image search. The main idea is to decompose the distance measure between variable-length descriptors into aligned component-to-component matching problems independently, and build multiple hash tables for the visual word components. Given a query, its candidate neighbors are found by using the query binary sub-vectors as indices into their corresponding hash tables. In particular, a bit selection based on conditional mutual information maximization is proposed to reduce the dimensionality of visual word components, which provides a light storage of indices and balances the retrieval accuracy and search cost. Extensive experiments on benchmark datasets show that our approach is 20~25 times faster than linear search, without any noticeable retrieval performance loss. Zhe Wang 0019, Ling-Yu Duan, Jie Lin 0001, Tiejun Huang 0001, Wen Gao 0001, Miroslaw Bober |
ICIP | 2 |
| 2014 | Interactive ads recommendation with contextual search on product topic space
Jinqiao Wang, Bo Wang 0011, Ling-Yu Duan, Qi Tian 0001, Hanqing Lu |
Multim. Tools Appl. | 3 |
| 2014 | Mining Compact Bag-of-Patterns for Low Bit Rate Mobile Visual SearchabstractVisual patterns, i.e., high-order combinations of visual words, contributes to a discriminative abstraction of the high-dimensional bag-of-words image representation. However, the existing visual patterns are built upon the 2D photographic concurrences of visual words, which is ill-posed comparing with their real-world 3D concurrences, since the words from different objects or different depth might be incorrectly bound into an identical pattern. On the other hand, designing compact descriptors from the mined patterns is left open. To address both issues, in this paper, we propose a novel compact bag-of-patterns (CBoPs) descriptor with an application to low bit rate mobile landmark search. First, to overcome the ill-posed 2D photographic configuration, we build up a 3D point cloud from the reference images of each landmark, therefore more accurate pattern candidates can be extracted from the 3D concurrences of visual words. A novel gravity distance metric is then proposed to mine discriminative visual patterns. Second, we come up with compact image description by introducing a CBoPs descriptor. CBoP is figured out by sparse coding over the mined visual patterns, which maximally reconstructs the original bag-of-words histogram with a minimum coding length. We developed a low bit rate mobile landmark search prototype, in which CBoP descriptor is directly extracted and sent from the mobile end to reduce the query delivery latency. The CBoP performance is quantized in several large-scale benchmarks with comparisons to the state-of-the-art compact descriptors, topic features, and hashing descriptors. We have reported comparable accuracy to the million-scale bag-of-words histogram over the million scale visual words, with high descriptor compression rate (approximately 100-bits) than the state-of-the-art bag-of-words compression scheme. Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Tiejun Huang 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2014 | Spatiotemporal Grid Flow for Video RetargetingabstractVideo retargeting is a useful technique to adapt a video to a desired display resolution. It aims to preserve the information contained in the original video and the shapes of salient objects while maintaining the temporal coherence of contents in the video. Existing video retargeting schemes achieve temporal coherence via constraining each region/pixel to be deformed consistently with its corresponding region/pixel in neighboring frames. However, these methods often distort the shapes of salient objects, since they do not ensure the content consistency for regions/pixels constrained to be coherently deformed along time axis. In this paper, we propose a video retargeting scheme to simultaneously meet the two requirements. Our method first segments a video clip into spatiotemporal grids called grid flows, where the consistency of the content associated with a grid flow is maintained while retargeting the grid flow. After that, due to the coarse granularity of grid, there still may exist content inconsistency in some grid flows. We exploit the temporal redundancy in a grid flow to avoid that the grids with inconsistent content be incorrectly constrained to be coherently deformed. In particular, we use grid flows to select a set of key-frames which summarize a video clip, and resize subgrid-flows in these key-frames. We then resize the remaining nonkey-frames by simply interpolating their grid contents from the two nearest retargeted key-frames. With the key-frame-based scheme, we only need to solve a small-scale quadratic programming problem to resize subgrid-flows and perform grid interpolation, leading to low computation and memory costs. The experimental results demonstrate the superior performance of our scheme. Bing Li 0024, Ling-Yu Duan, Jinqiao Wang, Rongrong Ji, Chia-Wen Lin, Wen Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2014 | Towards Mobile Document Image Retrieval for Digital LibraryabstractWith the proliferation of mobile devices, recent years have witnessed an emerging potential to integrate mobile visual search techniques into digital library. Such a mobile application scenario in digital library has posed significant and unique challenges in document image search. The mobile photograph makes it tough to extract discriminative features from the landmark regions of documents, like line drawings, as well as text layouts. In addition, both search scalability and query delivery latency remain challenging issues in mobile document search. The former relies on an effective yet memory-light indexing structure to accomplish fast online search, while the latter puts a bit budget constraint of query images over the wireless link. In this paper, we propose a novel mobile document image retrieval framework, consisting of a robust Local Inner-distance Shape Context (LISC) descriptor of line drawings, a Hamming distance KD-Tree for scalable and memory-light document indexing, as well as a JBIG2 based query compression scheme, together with a Retinex based enhancement and an OTSU based binarization, to reduce the latency of delivering query while maintaining query quality in terms of search performance. We have extensively validated the key techniques in this framework by quantitative comparison to alternative approaches. Ling-Yu Duan, Rongrong Ji, Tiejun Huang 0001, Wen Gao 0001 |
IEEE Trans. Multim. | 1 |
| 2013 | On the interoperability of local descriptors compressionabstractThere are a number of component technologies that are useful for visual search, including format of visual descriptors, descriptor extraction process, as well as indexing, and matching algorithms. As a minimum, the format of descriptors as well as parts of their extraction process should be defined to ensure interoperability. In this paper, we study the problem of interoperability among compressed local descriptors at different bit-rates; that is, allowing effective and efficient comparison of compact descriptors, which is fundamentally important to mobile visual search applications. We propose to combine feature transform and multi-stage vector quantization to implement the interoperability of compact local descriptors. First, an orthogonal transform (e.g. Principle component analysis, PCA) is employed to eliminate the correlation between local feature dimensions, which improves the performance of compressed domain descriptor matching with the well-aligned distance computing of sorted important features in transform space. Second, a multi-stage vector quantization (MSVQ) is applied to generate compact codes for local descriptors. At light quantization tables, MSVQ takes advantage of the transform domain features to properly allocate different budgets to each group of transformed feature dimensions, respectively. The interoperability between compressed descriptors at different bit rates can be achieved by the descriptors' fast matching in the orthogonal feature space. In other words, descriptor decoding into the original feature space (SIFT space) is unnecessary, as the distance can be calculated by pre-computed lookup tables. In particular, such efficient matching in transform domain is significant for large-scale visual search. Over a set of benchmark datasets, we have reported superior performance over state-of-the-arts. Jie Chen 0006, Ling-Yu Duan, Jie Lin 0001, Rongrong Ji, Tiejun Huang 0001, Wen Gao 0001 |
ICASSP | 2 |
| 2013 | Robust fisher codes for large scale image retrievalabstractFisher vectors (FV) have shown great advantages in large scale visual search. However, traditional FV suffers from noisy local descriptors, which may deteriorate the FV discriminative power. In this paper, we propose a robust Fisher vectors (RFV). To fulfill fast search and light storage over a large scale image dataset, we employ a simple binarization method to compress RFV to generate compact robust Fisher codes (RFC). Extensive comparison experiments on benchmark datasets have shown that both RFV and RFC outperforms the state-of-the-art performance. The scalability of RFC has been validated on a dataset of over 1 million images as well. Jie Lin 0001, Ling-Yu Duan, Tiejun Huang 0001, Wen Gao 0001 |
ICASSP | 2 |
| 2013 | A novel pair-wise image matching strategy with compact descriptorsabstractIn this paper, we address the problem of pair-wise image matching which determines whether two images depict the same objects or scenes. SIFT-like local descriptor-based matching is the most widely adopted method for this purpose and has achieved the state-of-the-art performance. However, local descriptor-based methods usually fail when an image pair contains multiple similar local regions. This problem becomes more serious when coming to limited computational and storage resources. Although global descriptors, e.g., Fisher Vectors, can solve this issue, it is difficult for global descriptors to distinguish images containing different objects of the same class. Therefore, we propose a novel strategy to integrate local and global descriptors for better matching accuracy. To further fulfill the efficiency requirement of applications, we combine dimension reduction and product quantization to obtain compact descriptors and speed up the matching process with pre-computed lookup tables. Extensive comparisons to the state-of-the-art methods demonstrate our advantages in both matching accuracy and efficiency. Ling-Yu Duan, Jie Lin 0001, Tiejun Huang 0001 |
ICIP | 2 |
| 2013 | Compact descriptors for mobile visual search and MPEG CDVS standardizationabstractIn this paper, we present the state-of-the-art compact descriptors for mobile visual search. In particular, we introduce our MPEG contributions in global descriptor aggregation and local descriptor compression, which have been adopted by the ongoing MPEG standardization of compact descriptor for visual search (CDVS). Standardization progress will be introduced. Other issues including visual object databases and MPEG CDVS impact on visual search industry will be discussed as well. Ling-Yu Duan, Feng Gao 0014, Jie Chen 0006, Jie Lin 0001, Tiejun Huang 0001 |
ISCAS | 1 |
| 2013 | Mobile media communication, processing, and analysis: A review of recent advancesabstractIn this paper, we review recent advances in mobile media communication, processing, and analysis. To identify the opportunities and challenges in fast growing mobile media computing, we discuss several emerging topics including mobile visual search, retargeting, mobile video streaming, and cloud based mobile media computing. According to the infrastructure of mobile devices vs. servers, we come up with essential concerns in mobile media computing such as wireless bandwidth consumption, mobile energy saving, media adaptation for better quality of services, the computational load shift from mobiles to servers, etc. With booming mobile Apps on diverse media consumption, it is envisioned that mobile media research and development is bringing about significant achievements in traditional topics of communication, processing, and analytics. Wen Gao 0001, Ling-Yu Duan, Jun Sun 0007, Junsong Yuan 0001, Yonggang Wen 0001, Yap-Peng Tan, Jianfei Cai 0001, Alex Chichung Kot |
ISCAS | 2 |
| 2013 | An Error Resilient Depth Map Coding Scheme Using Adaptive Wyner-Ziv Frame
Xiangkai Liu, Qiang Peng, Xiao Wu 0001, Lei Zhang 0006, Ling-Yu Duan |
MMM (2) | 6 |
| 2013 | A local shape descriptor for mobile linedrawing retrievalabstractComing with the rapid spread of Intelligent terminals with camera, mobile visual search techniques have undergone a revolution, where visual information can be easily browsed and retrieved upon simply capturing a query photo. However, most existing work targets at compact description of natural scene image statistics, while dealing with line drawing images retains an open problem. This paper presents a unified framework of line drawing problems in mobile visual search. We propose a compact description of line drawing image named Local Inner-Distance Shape Context (LISC) which is robust to the distortion and occlusion and enjoys scale and rotation invariance. Together with an innovative compression scheme using JBIG2 to reduce query delivery latency, our framework works well on both a self-built dataset and MPEG- 7 CE Shape-1 dataset. Promising results on both datasets show significant improvement over state-of-the-art algorithms. Yucong Xuan, Ling-Yu Duan, Tiejun Huang 0001 |
VCIP | 2 |
| 2013 | Learning from mobile contexts to minimize the mobile location search latency
Ling-Yu Duan, Rongrong Ji, Jie Chen 0006, Hongxun Yao, Tiejun Huang 0001, Wen Gao 0001 |
Signal Process. Image Commun. | 1 |
| 2013 | Estimating Visual Saliency Through Single Image OptimizationabstractThis letter presents a novel approach for visual saliency estimation through single image optimization. Instead of directly mapping visual features to saliency values with a unified model, we treat regional saliency values as the optimization objective on each single image. By using a quadratic programming framework, our approach can adaptively optimize the regional saliency values on each specific image to simultaneously meet multiple saliency hypotheses on visual rarity, center-bias and mutual correlation. Experimental results show that our approach can outperform 14 state-of-the-art approaches on a public image benchmark. Jia Li 0003, Yonghong Tian 0001, Ling-Yu Duan, Tiejun Huang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2013 | Learning to Distribute Vocabulary Indexing for Scalable Visual SearchabstractIn recent years, there is an ever-increasing research focus on Bag-of-Words based near duplicate visual search paradigm with inverted indexing. One fundamental yet unexploited challenge is how to maintain the large indexing structures within a single server subject to its memory constraint, which is extremely hard to scale up to millions or even billions of images. In this paper, we propose to parallelize the near duplicate visual search architecture to index millions of images over multiple servers, including the distribution of both visual vocabulary and the corresponding indexing structure. We optimize the distribution of vocabulary indexing from a machine learning perspective, which provides a “memory light” search paradigm that leverages the computational power across multiple servers to reduce the search latency. Especially, our solution addresses two essential issues: “What to distribute” and “How to distribute”. “What to distribute” is addressed by a “lossy” vocabulary Boosting, which discards both frequent and indiscriminating words prior to distribution. “How to distribute” is addressed by learning an optimal distribution function, which maximizes the uniformity of assigning the words of a given query to multiple servers. We validate the distributed vocabulary indexing scheme in a real world location search system over 10 million landmark images. Comparing to the state-of-the-art alternatives of single-server search,,and distributed search, our scheme has yielded a significant gain of about 200% speedup at comparable precision by distributing only 5% words. We also report excellent robustness even when partial servers crash. Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Lexing Xie, Hongxun Yao, Wen Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2012 | Towards compact topical descriptorsabstractWe introduce a Compact Topical Descriptor to learn a compact yet discriminative image signature from the reference image corpus. This descriptor is deployed over the well used bag-of-words image histogram, with two merits over the traditional topical features: First, we propose to directly control the topical sparsity to achieve the descriptor compactness. Second, we ensure the descriptor discriminability by minimizing the bag-of-words reconstruction errors during the topical histogram encoding. To this end, we have a generative viewpoint of the topical feature extraction, which is estimated as a sparse MAP estimation over the original bag-of-words. We learn such estimation by a bi-convex optimization, iterating between both hierarchical sparse coding from words to topical histograms and dictionary learning of the corresponding word-to-topic transform. Especially, supervised labels such as image ranking list can be also incorporated into our descriptor learning paradigm. We quantize our performance in both Im-ageNet 10K and NUS-WIDE, with comparisons to bag-of-words, LDA, miniBoF, and Aggregated Local Descriptors. In practice, we also implement our descriptor for a low bit rate mobile visual search application, i.e. sending compact descriptors instead of the image to reduce the query delivery latency. Our descriptor has significantly outperformed the state-of-the-art compact descriptors by quantitative evaluations over 10 million reference images. Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Wen Gao 0001 |
CVPR | 2 |
| 2012 | Pruning tree-structured vector quantizer towards low bit rate mobile visual searchabstractComing with the proliferation of mobile devices, mobile visual search emerges. One fundamental issue here is the query transmission latency, especially in a bandwidth constraint wireless link. Towards low bit rate retrieval, recent works have proposed to extract compact visual descriptors directly on the mobile end, where the vocabulary tree based bag-of-words representation has shown superior performance in producing compact descriptors [2][9]. However, the corresponding tree-structure vector quantizer is extremely large against a mobile end implementation. In this paper, we propose two alternatives to prune this tree structure based on the subtree discriminability analysis, where either information gain based or ranking based pruning are investigated. Furthermore, we have unveiled that the tree structure can be even discarded while retaining only the discriminative leaves together with their radii in practice. We evaluate our tree pruning on Android HTC Desire G7, with application to low bit rate mobile landmark search in a 10-million landmark photo collection, where over 10 scale memory reduction with almost identical search accuracy is reported. Jie Chen 0006, Ling-Yu Duan, Rongrong Ji, Wen Gao 0001 |
ICASSP | 2 |
| 2012 | Predicting the effectiveness of queries for visual searchabstractPoor retrieval performance significantly degenerates users' experience of visual search, especially in mobile search. Ideally, users would like to be alerted when bad queries are present, which helps eliminate latency as well as waste of bandwidth, especially in 3G wireless environment. In this paper, we propose a visual query performance prediction (v-QPP) approach to predict the retrieval effectiveness. We employ latent dirichlet allocation (LDA)to derive latent topics from image database. From the collection statistics, we model the query's specificity based on topics. High specificity helps a retrieval system to derive user's search intent exactly. Moreover, as low discriminative content is difficult to search in terms of distinguishing relevant images from irrelevant one, we propose a topics based inverse concept frequency (t-ICF) model to deal with specific queries but difficult to discriminate in the reference database. Comparison experiments over MPEG CDVS benchmarking datasets have shown our method significantly outperforms existing approaches in document retrieval. Bing Li 0024, Ling-Yu Duan, Rongrong Ji, Wen Gao 0001 |
ICASSP | 2 |
| 2012 | Learning multiple codebooks for low bit rate mobile visual searchabstractCompressing a query image's signature via vocabulary coding is an effective approach to low bit rate mobile visual search. State-of-the-art methods concentrate on offline learning a codebook from an initial large vocabulary. Over a large heterogeneous reference database, learning a single codebook may not suffice for maximally removing redundant codewords for vocabulary based compact descriptor. In this paper, we propose to learn multiple codebooks (m-Codebooks) for extremely compressing image signatures. A query-specific codebook (q-Codebook) is online generated at both client and server sides by adaptively weighting the off-line learned multiple codebooks. The q-Codebook is subsequently employed to quantize the query image for producing compact, discriminative, and scalable descriptors. As q-Codebook may be simultaneously generated at both sides, without transmitting the entire vocabulary, only small overhead (e.g. codebook ID and codeword 0/1 index) is incurred to reconstruct the query signature at the server end. To fulfill m-Codebooks and q-Codebook, we adopt a Bi-layer Sparse Coding method to learn the sparse relationships of codewords vs. codebooks as well as codebooks vs. query images via l1 regularization. Experiments on benchmarking datasets have demonstrated the extremely small descriptor's supervior performance in image retrieval. Jie Lin 0001, Ling-Yu Duan, Jie Chen 0006, Rongrong Ji, Siwei Luo, Wen Gao 0001 |
ICASSP | 2 |
| 2012 | PQ-WGLOH: A bit-rate scalable local feature descriptorabstractIn this paper, we propose a compact yet discriminative local descriptor which tackles the wireless query transmission latency in mobile visual search. The descriptor captures gradient statistics of canonical patches over a log-polar location grid whose parameters are optimized using training samples. We quantize the resulting descriptor using product quantization. The descriptor achieves about 95% bits reduction compared with 128-Byte SIFT and allows adaptation of descriptor lengths to support user required performance. Moreover, accurate matching of descriptors with low complexity is allowed within several table lookup operations. We perform a comprehensive comparison with SIFT, GLOH and CHoG in the context of image retrieval, image matching and object localization. We achieve competing matching and retrieval performance with SIFT, GLOH with much fewer bits. In particular, the descriptor outperforms CHoG at the same bits on eight data sets contributed to MPEG Compact Descriptor for Visual Search(CDVS) Standardization. Chunyu Wang 0001, Ling-Yu Duan, Yizhou Wang 0001, Wen Gao 0001 |
ICASSP | 2 |
| 2012 | Multi-stage vector quantization towards low bit rate visual searchabstractWhile much progress has been made in mobile visual search, user experiences still relate to the query transmission latency, especially over a bandwidth-constrained wireless link. Low bit rate visual search paradigm has been well advocated in both academic and industrial endeavors, which directly extracts and sends compact visual descriptor(s) rather than sending a query image. Recent advances in compact descriptor design have advocated the use of compressed bag-of-words histogram, which has shown superior performance over other alternatives. However, existing works focus on descriptor compactness, regardless of time cost and memory requirements on the extraction pipeline, which in turn is crucial for the mobile end development. In this paper, we investigate the problem of designing a memory-light descriptor extraction scheme based upon the so-called multi-stage vector quantization. Our scheme starts by quantizing local patches with a small codebook, and the resulting quantization residual is subsequently compensated by a product quantizer. The design of both quantizers are based upon improving PSNR, which would drop a lot through quantization. PSNR is quantitatively shown to be highly correlated with retrieval and matching accuracy. Extensive evaluation on MPEG Compact Descriptor for Visual Search (CDVS) dataset, has reported superior performance over the state-of-the-art. Jie Chen 0006, Ling-Yu Duan, Rongrong Ji, Zhe Wang 0019 |
ICIP | 2 |
| 2012 | Weakly supervised topic grouping of YouTube search resultsabstractRecent years have witnessed an explosive growth of user contributed videos on websites like YouTube and Metacafe, which usually provide a query-by-keyword functionality to facilitate the user browsing. For a given query, the returned videos typically contain multiple topics that are mixed up to duplicate the user browsing. Therefore, their diversification and grouping are highly demanded to improve the user experiences. However, the tagging and content qualities of user contributed videos are uncontrolled against their precise grouping. In this paper, we present a weakly supervised topic grouping paradigm to diversify the returned videos of a given keyword query. Our grouping is based on the bag-of-words visual signature quantized over the spatiotemporal STIP descriptor [1] extracted from each returned video. First, we adopt a min-Hashing based visual similarity in combination of the tagging similarity to group the returned videos. Based on the initial grouping configurations, we mine the co-occurred discriminative sub-signatures, based on which we iteratively refine the first step. Such iteration well handles the noise in visual content and tagging, since neither of which is fully trusted during the grouping. We validate our schemes on over 2,000 video clips crawled from a set of YouTube keyword query results. Comparing to alternative approaches, our scheme has shown superior robustness and precision. Liujuan Cao, Rongrong Ji, Wei Liu 0005, Yue Gao 0002, Ling-Yu Duan, Chaoguang Men |
ICIP | 5 |
| 2012 | Learning sparse tag patterns for social image classificationabstractUser-generated tags associated with images from social media (e.g., Flickr) provide valuable textual resources for image classification. However, the noisy and huge tag vocabulary heavily degrades the effectiveness and efficiency of state-of-the-art image classification methods that exploited auxiliary web data. To alleviate the problem, we introduce a Sparse Tag Patterns (STP) model to discover sparsity constrained co-occurrence tag patterns from large scale user contributed tags among social data. To fulfill the compactness and discriminability, we formulate STP as a problem of minimizing a quadratic loss function regularized by the bi-layer l1norm. We treat the learned STP as alternative intermediate semantic image feature and verify its superiority within a search-based image classification framework. Experiments on 240K social images associated with millions of tags have demonstrated encouraging performance of the proposed method compared to the state-of-the-art. Jie Lin 0001, Ling-Yu Duan, Junsong Yuan 0001, Qingyong Li, Siwei Luo |
ICIP | 2 |
| 2012 | Social Image Tagging by Mining Sparse Tag Patterns from Auxiliary DataabstractUser-given tags associated with social images from photosharing websites (e.g., Flickr) are valuable auxiliary resources for the image tagging task. However, social images often suffer from noisy and incomplete tags, heavily degrading the effectiveness of previous image tagging approaches. To alleviate the problem, we introduce a Sparse Tag Patterns (STP) model to discover noiseless and complementary cooccurrence tag patterns from large scale user contributed tags among auxiliary web data. To fulfill the compactness and discriminability, we formulate the STP model as a problem of minimizing quadratic loss function regularized by bi-layer ℓ1norm. We treat the learned STP as a universal knowledge base and verify its superiority within a data-driven image tagging framework. Experimental results over 1 million auxiliary data demonstrate superior performance of the proposed method compared to the state-of-the-art. Jie Lin 0001, Junsong Yuan 0001, Ling-Yu Duan, Siwei Luo, Wen Gao 0001 |
ICME | 3 |
| 2012 | Motion Based Perceptual Distortion and Rate Optimization for Video CodingabstractMost conventional distortion metrics regard a video frame as a static image, and seldom exploit using the motion information of video frames in succession. Moreover, these methods usually calculate the visual distortion based on the independent spatial pixels. Recently, many researches show that the way people perceive the video signals is similar to the way filters process signals in the frequency domain. Therefore, in order to achieve better visual quality, we introduce a novel distortion measurement into the video coding system, which is consistent with human visual perception, and establish a perception-based rate-distortion optimization model. In this paper, we adopt Gabor filter family to decompose the video signals into frequency domain, and combine the video motion information to measure the perceptual distortion. We call it Motion tuned Distortion metric For Video coding (MDFV). After that we set up an MDFV based rate-distortion optimization model to select the best encoding mode. The experimental results show that the proposed approach is effective. Xi Wang 0014, Li Su 0003, Qingming Huang, Chunxi Liu, Ling-Yu Duan |
ICME | 5 |
| 2012 | Optimizing JPEG quantization table for low bit rate mobile visual searchabstractSmart phones is bringing about emerging potentials in mobile visual search. Extensive research efforts have been made in compact visual descriptors. However, directly extracting visual descriptors on a mobile device is computationally intensive and time consuming. Towards low bit rate visual search, we propose to deeply compress query images by learning a customized JPEG quantization table in the context of visual search. Distinct from traditional image compression, by incorporating pair-wise image matching precision into distortion measure, we optimize quantization table to seek a better trade-off between image compression rate and visual search performance. An evolutionary algorithm is employed to learn an optimal quantization table. Under MPEG CDVS evaluation framework, extensive evaluation has been done including image retrieval and pair-wise matching over 1 million database images. Experimental results have demonstrated that our optimized quantization table works much better than JPEG default one in terms of retrieval/matching performance vs. a set of different operating points. The proposed low bit rate solution may be easily deployed to smart phones without hardware support, as a useful complement to the ongoing MPEG CDVS standardization efforts. Ling-Yu Duan, Xiangkai Liu, Jie Chen 0006, Tiejun Huang 0001, Wen Gao 0001 |
VCIP | 1 |
| 2012 | Location Discriminative Vocabulary Coding for Mobile Landmark Search
Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Hongxun Yao, Junsong Yuan 0001, Yong Rui, Wen Gao 0001 |
Int. J. Comput. Vis. | 2 |
| 2012 | Group-Sensitive Multiple Kernel Learning for Object RecognitionabstractIn this paper, a group-sensitive multiple kernel learning (GS-MKL) method is proposed for object recognition to accommodate the intraclass diversity and the interclass correlation. By introducing the "group" between the object category and individual images as an intermediate representation, GS-MKL attempts to learn group-sensitive multikernel combinations together with the associated classifier. For each object category, the image corpus from the same category is partitioned into groups. Images with similar appearance are partitioned into the same group, which corresponds to the subcategory of the object category. Accordingly, intraclass diversity can be represented by the set of groups from the same category but with diverse appearances; interclass correlation can be represented by the correlation between groups from different categories. GS-MKL provides a tractable solution to adapt multikernel combination to local data distribution and to seek a tradeoff between capturing the diversity and keeping the invariance for each object category. Different from the simple hybrid grouping strategy that solves sample grouping and GS-MKL training independently, two sample grouping strategies are proposed to integrate sample grouping and GS-MKL training. The first one is a looping hybrid grouping method, where a global kernel clustering method and GS-MKL interact with each other by sharing group-sensitive multikernel combination. The second one is a dynamic divisive grouping method, where a hierarchical kernel-based grouping process interacts with GS-MKL. Experimental results show that performance of GS-MKL does not significantly vary with different grouping strategies, but the looping hybrid grouping method produces slightly better results. On four challenging data sets, our proposed method has achieved encouraging performance comparable to the state-of-the-art and outperformed several existing MKL methods. Yonghong Tian 0001, Ling-Yu Duan, Tiejun Huang 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2012 | A Generic Approach for Systematic Analysis of Sports VideosabstractVarious innovative and original works have been applied and proposed in the field of sports video analysis. However, individual works have focused on sophisticated methodologies with particular sport types and there has been a lack of scalable and holistic frameworks in this field. This article proposes a solution and presents a systematic and generic approach which is experimented on a relatively large-scale sports consortia. The system aims at the event detection scenario of an input video with an orderly sequential process. Initially, domain knowledge-independent local descriptors are extracted homogeneously from the input video sequence. Then the video representation is created by adopting a bag-of-visual-words (BoW) model. The video’s genre is first identified by applying the k-nearest neighbor (k-NN) classifiers on the initially obtained video representation, and various dissimilarity measures are assessed and evaluated analytically. Subsequently, an unsupervised probabilistic latent semantic analysis (PLSA)-based approach is employed at the same histogram-based video representation, characterizing each frame of video sequence into one of four view groups, namely closed-up-view, mid-view, long-view, and outer-field-view. Finally, a hidden conditional random field (HCRF) structured prediction model is utilized for interesting event detection. From experimental results, k-NN classifier using KL-divergence measurement demonstrates the best accuracy at 82.16% for genre categorization. Supervised SVM and unsupervised PLSA have average classification accuracies at 82.86% and 68.13%, respectively. The HCRF model achieves 92.31% accuracy using the unsupervised PLSA based label input, which is comparable with the supervised SVM based input at an accuracy of 93.08%. In general, such a systematic approach can be widely applied in processing massive videos generically. Ning Zhang 0023, Ling-Yu Duan, Lingfang Li, Qingming Huang, Wen Gao 0001, Ling Guan |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2011 | Sorting local descriptors for lowbit rate mobile visual searchabstractState-of-the-art mobile visual search systems put emphasis on developing compact visual descriptors, which enables low bit rate wireless transmission instead of delivering an entire query image. In this paper, we address the orderless nature of the transmission set of query descriptors . We propose to adapt the orders of local descriptors in transmission, which subsequently yields more consistent statistic distributions in each feature dimension towards more efficient residual coding based compression. Our scheme further enables lossy sorting by an adaptive quantization strategy within each feature dimension, which largely improves the compression rates of the residual coding in each dimension. We show that the performance degeneration of such lossy sorting is acceptable in our mobile landmark search applications. Our approach's effectiveness and efficiency is demonstrated via extensive experimental comparisons to state-of-the art works in both mobile visual descriptors and compact image signatures. Jie Chen 0006, Ling-Yu Duan, Rongrong Ji, Hongxun Yao, Wen Gao 0001 |
ICASSP | 2 |
| 2011 | A lowbit rate vocabulary coding scheme for mobile landmark searchabstractWe present a low bit rate vocabulary coding scheme in the context of mobile landmark search. Our scheme exploits location cues to boost a compact subset of visual vocabulary, which is discriminative for visual search and incurs low bit rate query for efficient upstream wireless transmission. To validate the coding scheme, we have developed mobile landmark search prototype systems within typical areas including Beijing, New York City, Lhasa, Singapore, and Florence. Our system maintains a single vocabulary in a mobile device, which can be efficiently adapted with the location information of city-scale mobile users. Thus multiple downloading of large vocabulary is completely avoided for normal city tourists. In landmark search domain, we have reported superior performance over the state-of-the art works in compact image descriptors or signatures. Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Hongxun Yao, Wen Gao 0001 |
ICASSP | 2 |
| 2011 | When codeword frequency meets geographical locationabstractWhen codeword frequency meets geographical location in landmark search applications, is it still discriminative for the search procedure. In this paper, we give a systematic investigation about how geographical location affects the effectiveness of codeword frequency. We explain why the standard IDF in the BoW models is less effective in location related search applications [11][12]. Consequently, we propose a “location discriminative codeword frequency” strategy to introduce the location context into the codeword discriminability measurement. This new codeword frequency is calculated in each geographical region, for which a spectral clustering scheme is proposed to partition the geographical map of each city into distinct regions. Extensive comparisons over the standard codeword frequency in state-of-the-art landmark search systems [1][1] demonstrates our approach's effectiveness. Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Hongxun Yao, Wen Gao 0001 |
ICASSP | 2 |
| 2011 | Generating vocabulary for global feature representation towards commerce image retrievalabstractThis paper studies the problem of retrieving images by color, texture and shape in the context of visual assisted product recommendation in E-commerce sites. Different from general CBIR applications, commerce image retrieval puts more emphasis on outlier-free ranking (top N) to gain perfect user experience. We suggest to extend the bag-of-words (BoW) model to global feature characterization rather than commonly used histogram based low-level feature representation. Although BoW is a common practice in object recognition, we argue generating feature vocabulary is useful to address the global feature characterization that could be elegantly adapted to domain specific commerce image search. The representation is compact and discriminative, which may adapt with individual websites. Quantitative as well as subjective evaluation demonstrates the functionality of the proposed method. In practice, the vocabulary based global features greatly reduce outliers in top rank images, so that desirable user experience can be obtained in E-Commerce applications. Ling-Yu Duan, Chunyu Wang 0001, Tiejun Huang 0001, Wen Gao 0001 |
ICIP | 2 |
| 2011 | PKUBench: A context rich mobile visual search benchmarkabstractWhile there are ever growing focuses on mobile visual search in recent years, a comprehensive benchmark database with rich context information (such as GPS) for fair evaluation among different strategies is still missing. This paper introduces a PKUBench benchmark for the quantitative evaluations of mobile visual search with the support of GPS. It contains 13,179 images organized into 198 distinct landmark locations within the Peking University campus. Each location is captured with multiple shot sizes and viewing angles, using both digital cameras and phone cameras, each photo being tagged with rich contextual information in the mobile scenario. Moreover, this benchmark studies typical quality degeneration scenarios in mobile photographing, including variable resolutions, blurring, lighting changes, occlusions, as well as various viewing angles. Together with this benchmark, we provide the bag-of-visual-words search baselines involving contextual information refinement. Finally, distractor images are further introduced to evaluate the robustness of visual search methods in this database. Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Tiejun Huang 0001, Hongxun Yao, Wen Gao 0001 |
ICIP | 2 |
| 2011 | Learning the trip suggestion from landmark photos on the webabstractIn this paper, we introduce a novel touristic trip suggestion system to facilitate the traveling of mobile users in a given city. Given the current user location and his touristic destination, our system can suggest a shortest trip path that visits as many popular landmarks as possible. To this end, we collect geographical tagged photos from Flickr [1] and Panoramio [2] photo sharing websites. Then a geographical graph is constructed by modeling photos as vertices and their geographical and visual closenesses as connection strengths. In this graph, we mine a dominant subgraph by quantizing nearby and visually duplicated vertices, and then trimming unpopular subgraphs. Such dominant subgraph only retains the popular landmarks from the consensus of travelers in this city. In online suggestion, we map the current user location and the target location to the nearest vertices in this subgraph, based on which an optimal trip is suggested through a shortest path search. We have quantitatively validated our system in typical areas including Beijing and New York City, with quantitative comparisons to alternative approaches. Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Hongxun Yao, Tiejun Huang 0001, Wen Gao 0001 |
ICIP | 2 |
| 2011 | Fast retargeting with adaptive grid optimizationabstractEffective and efficient retargeting techniques may enrich users' browsing experiences in mobile devices. Existing mesh-based retargeting solutions put less efforts in making well-tuned meshes. In this paper, we propose a novel adaptive grid based optimization method to retarget an image. First, we present an entropy based measure to guide the grid construction. Then we employ the quadtree structure to adjust the grid granularity adaptively. Furthermore, to reduce the inappropriate deformation from inconsistent importance assignment, we build a global optimization model to alleviate serious shape deformation in retargeting. Comparison experiments show our method's superiority over the state-of-the-art approaches. Bing Li 0024, Jinqiao Wang, Ling-Yu Duan, Wen Gao 0001 |
ICME | 4 |
| 2011 | Learning Compact Visual Descriptor for Low Bit Rate Mobile Landmark Search
Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Hongxun Yao, Tiejun Huang 0001, Wen Gao 0001 |
IJCAI | 2 |
| 2011 | Towards low bit rate mobile visual search with multiple-channel codingabstractIn this paper, we propose a multiple-channel coding scheme to extract compact visual descriptors for low bit rate mobile visual search. Different from previous visual search scenarios that send the query image, we make use of the ever growing mobile computational capability to directly extract compact visual descriptors at the mobile end. Meanwhile, stepping forward from the state-of-the-art compact descriptor extractions, we exploit the rich contextual cues at the mobile end (such as GPS tags for mobile visual search and 2D barcodes or RFID tags for mobile product search), together with the visual statistics at the reference database, to learn multiple coding channels. Therefore, we describe the query with one of many forms of high-dimensional visual signature, which is subsequently mapped to one or more channels and compressed. The compression function within each channel is learnt based on a novel robust PCA scheme, with specific consideration to preserve the retrieval ranking capability of the original signature. We have deployed our scheme on both iPhone4 and HTC DESIRE 7 to search ten million landmark images in a low bit rate setting. Quantitative comparisons to the state-of-the-arts demonstrate our significant advantages in descriptor compactness (with orders of magnitudes improvement) and retrieval mAP in mobile landmark, product, and CD/book cover search. Rongrong Ji, Ling-Yu Duan, Jie Chen 0006, Hongxun Yao, Yong Rui, Shih-Fu Chang, Wen Gao 0001 |
ACM Multimedia | 2 |
| 2011 | Grid-Based Retargeting with Transformation Consistency Smoothing
Bing Li 0024, Ling-Yu Duan, Jinqiao Wang, Jie Chen 0006, Rongrong Ji, Wen Gao 0001 |
MMM (2) | 2 |
| 2010 | ESUR: A system for Events detection in SURveillance videoabstractIn this paper, we present our eSur (Event detection system on SURveillance video) system, which is derived from TRECVID'09 surveillance tasks. Currently, eSur attempts to detect two categories of events: 1) single-actor events (i.e., PersonRuns and ElevatorNoEntry) irrespective of any interaction between individuals, and 2) pair-activity events (i.e., PeopleMeet, PeopleSplitUp, and Embrace) involves more than one individual. eSur consists of three major stages, i.e., preprocessing, event classification, and post-processing. The preprocessing involves view classification, background subtraction, head-shoulder detection, human body detection and object tracking. Event classification fuses One-vs.-All SVM and rule-based classifiers to identify single-actor and pair-activity events in an ensemble way. To reduce false alarms, we introduce prior knowledge into the post-processing, and in particular, we apply a so-called event merging process over TRECVID dataset. Extensive experiments have been performed over TRECVid'08 and '09 ED data corpus involving in total 144 hours surveillance video of London Gatwick airport. According to the TRECVid-ED formal evaluation, our prototype has yielded fairly promising results over TRECVid'09 dataset, with top Act.DCR of 1.023, 1.025, 1.02, and 0.334 for PeopleMeet, PeopleSplitUp, Embrace, and ElevatorNoEntry, respectively. Yaowei Wang 0001, Yonghong Tian 0001, Ling-Yu Duan, Zhipeng Hu, Guochen Jia |
ICIP | 3 |
| 2010 | Interactive Web Video Advertising with Context Analysis and SearchabstractOnline media services and electronic commerce are booming recently. Previous studies have been devoted to contextual advertising, but few work deals with interactive web advertising. In this paper, we propose to put users in the loop of collecting contextual ad information with an interaction process, establishing semantic ad links across media platforms. Given an ad video, the key frames with explicit product information are located, which allow users to click favorite key frames for searching ads interactively. A three-stage contextual search is applied to find relevant products or services from web pages, i.e., searching visually similar product images on shopping websites, ranking product tags by text aggregation, and re-search textual items consisting of semantic meaningful tags to make a recommendation. In addition, users can choose automatically suggested keywords to reflect their intentions. Subjective evaluation has demonstrated the effectiveness of the proposed approach to interactive video advertising over the Web. Bo Wang 0011, Jinqiao Wang, Ling-Yu Duan, Qi Tian 0001, Hanqing Lu, Wen Gao 0001 |
ICPR | 3 |
| 2010 | Saliency detection based on 2D log-gabor wavelets and center biasabstractVisual saliency can be a useful tool for image content analysis such as automatic image cropping and image compression. In existing methods on visual saliency detection, most of them are related to the model of receptive field. In this paper, we propose a bottom-up model which introduces 2D Log-Gabor wavelets for saliency detection. Compared with the traditional model of receptive field, the 2D Log-Gabor wavelets can better simulate the biological characteristics of the simple cortical cell in the receptive filed. Moreover, we also incorporate the influence of center bias into our model, which is a common phenomenon that directs visual attention to the center of images in natural scenes. Experimental results show that our approach outperforms three state-of-the-art approaches remarkably. Jia Li 0003, Tiejun Huang 0001, Yonghong Tian 0001, Ling-Yu Duan, Guochen Jia |
ACM Multimedia | 5 |
| 2010 | AdVR: Linking Ad Video with Products or Service
Shi Chen 0008, Jinqiao Wang, Bo Wang 0011, Ling-Yu Duan, Qi Tian 0001, Hanqing Lu |
MMM | 4 |
| 2010 | Sequence Multi-Labeling: A Unified Video Annotation Scheme With Spatial and Temporal ContextabstractAutomatic video annotation is a challenging yet important problem for content-based video indexing and retrieval. In most existing works, annotation is formulated as a multi-labeling problem over individual shots. However, video is by nature informative in spatial and temporal context of semantic concepts. In this paper, we formulate video annotation as a sequence multi-labeling (SML) problem over a shot sequence. Different from many video annotation paradigms working on individual shots, SML aims to predict a multi-label sequence for consecutive shots in a global optimization manner by incorporating spatial and temporal context into a unified learning framework. A novel discriminative method, called sequence multi-label support vector machine (SVMSML), is accordingly proposed to infer the multi-label sequence for a given shot sequence. In SVMSML, a joint kernel is employed to model the feature-level and concept-level context relationships (i.e., the dependencies of concepts on the low-level features, spatial and temporal correlations of concepts). A multiple-kernel learning (MKL) algorithm is developed to optimize the kernel weights of the joint kernel as well as the SML score function. To efficiently search the desirable multi-label sequence over the large output space in both training and test phases, we adopt an approximate method to maximize the energy of a binary Markov random field (BMRF). Extensive experiments on TRECVID'05 and TRECVID'07 datasets have shown that our proposed SVMSMLgains superior performance over the state-of-the-art. Yuanning Li, Yonghong Tian 0001, Ling-Yu Duan, Tiejun Huang 0001, Wen Gao 0001 |
IEEE Trans. Multim. | 3 |
| 2009 | Group-sensitive multiple kernel learning for object categorizationabstractIn this paper, we propose a group-sensitive multiple kernel learning (GS-MKL) method to accommodate the intra-class diversity and the inter-class correlation for object categorization. By introducing an intermediate representation “group” between images and object categories, GS-MKL attempts to find appropriate kernel combination for each group to get a finer depiction of object categories. For each category, images within a group share a set of kernel weights while images from different groups may employ distinct sets of kernel weights. In GS-MKL, such group-sensitive kernel combinations together with the multi-kernels based classifier are optimized in a joint manner to seek a trade-off between capturing the diversity and keeping the invariance for each category. Extensive experiments show that our proposed GS-MKL method has achieved encouraging performance over three challenging datasets. Yuanning Li, Yonghong Tian 0001, Ling-Yu Duan, Wen Gao 0001 |
ICCV | 4 |
| 2009 | Linking video ADS with product or service information by web searchabstractWith the proliferation of online media services, video ads are pervasive across various platforms involving Internet services and interactive TV services. Existing research efforts such as Google AdSense and MSRA videosense/imagesense have been devoted to the less intrusive insertion of relevant textual or video ads in streams or Web pages through text/image/video content analysis whereas the inherent semantics of video ads is much less exploited. In this paper, we propose to link video ads with relevant product/service information across e-commerce Web sites or portals towards ad recommendation in a cross-media manner. Firstly, we carry out semantic analysis within ad videos in which frames marked with product images (FMPI) are extracted. Secondly, we link ad videos with relevant ads on the Web by utilizing FMPI to search visually similar product images (e.g. appearance or logo) and to collect their accompanying text (brand name, category, description, or other tags) over popular e-commerce Websites or portals such as EBay, Amazon, Taobao, etc. We search visually similar product images with local sensitive hashing (LSH) in a naive Bayes near neighbor classifier. Finally, we may recommend more relevant products/services for ad videos through ranking those matched product images and categorizing useful tags of top ranked ads from the Web. Preliminary experiments have been carried out to demonstrate the idea of linking ad videos with product/service information from the Web. Jinqiao Wang, Ling-Yu Duan, Bo Wang 0011, Shi Chen 0008, Jing Liu 0001, Hanqing Lu, Wen Gao 0001 |
ICME | 2 |
| 2009 | Multiple kernel active learning for image classificationabstractRecently, multiple kernel learning (MKL) methods have shown promising performance in image classification. As a sort of supervised learning, training MKL-based classifiers relies on selecting and annotating extensive dataset. In general, we have to manually label large amount of samples to achieve desirable MKL-based classifiers. Moreover, MKL also suffers a great computational cost on kernel computation and parameter optimization. In this paper, we propose a local adaptive active learning (LA-AL) method to reduce the labeling and computational cost by selecting the most informative training samples. LA-AL adopts a top-down (or global-local) strategy for locating and searching informative samples. Uncertain samples are first clustered into groups, and then informative samples are consequently selected via inter-group and intra-group competitions. Experiments over COREL-5K show that the proposed LA-AL method can significantly reduce the demand of sample labeling and have achieved the state-of-the-art performance. Yuanning Li, Yonghong Tian 0001, Ling-Yu Duan, Wen Gao 0001 |
ICME | 4 |
| 2009 | Automatic sports genre categorization and view-type classification over large-scale datasetabstractThis paper presents a framework with two automatic tasks targeting large-scale and low quality sports video archives collected from online video streams. The framework is based on the bag of visual-words model using speeded-up robust features (SURF). The first task is sports genre categorization based on hierarchical structure. Following on the second task which is based on automatically obtained genre, views are classified using support vector machines (SVMs). As a consequence, the views classification result can be used in video parsing and highlight extraction. As compared with state-of-the-art methods, our approach is fully automatic as well as domain knowledge free and thus provides a better extensibility. Furthermore, our dataset consists of 14 sport genres with 6850 minutes in total. Both sport genre categorization and view type classification have more than 80% accuracy rates, which validate this framework's robustness and potential in web-based applications. Lingfang Li, Ning Zhang 0023, Ling-Yu Duan, Qingming Huang, Ling Guan |
ACM Multimedia | 3 |
| 2009 | Consumer video retargeting: context assisted spatial-temporal grid optimizationabstractPervasive multimedia devices require accurate video retargeting, especially in connected consumer electronics platforms. In this paper, we present a context assisted spatialtemporal grid scheme for consumer video retargeting. First, we parse consumer videos from low-level features to highlevel visual concepts, combining visual attention into a more accurate importance description. Then, a semantic importance map is built up representing the spatial importance and temporal continuity, which is incorporated with a 3D rectilinear grid scaleplate to map frames to the target display, thereby keeping the aspect ratio of semantically salient objects as well as the perceptual coherency. Extensive evaluations were done on two popular video genres, sports and advertisements. The comparison with state-of-the-art approaches on both images and videos have demonstrated the advantages of the proposed approach. Jinqiao Wang, Ling-Yu Duan, Hanqing Lu |
ACM Multimedia | 3 |
| 2009 | Sports video retargetingabstractWith the proliferation of diverse multimedia terminals, the request for elegantly retargeting videos to different display devices is evident, especially in sports. This demonstration presents a Sports Video Retargeting(SVR) technique, that utilized domain based structure parsing to build a semantic importance map for video retargeting. The system enables flexible and coherent aspect-ratio change of the output sports videos with a spatial-temporal 3D rectilinear grid framework, which are free from significant loss of information or distortion on salient and important regions. Results in various sports type have shown that SVR is promising for content adaptation on mobile media. Jinqiao Wang, Ling-Yu Duan, Hanqing Lu |
ACM Multimedia | 3 |
| 2009 | A New Multiple Kernel Approach for Visual Concept Learning
Yuanning Li, Yonghong Tian 0001, Ling-Yu Duan, Wen Gao 0001 |
MMM | 4 |
| 2008 | Personalization of media and its attention service applicationsabstractA wealth of information creates a poverty of attention and a need to allocate that attention efficiently among the overabundance of information sources that might consume it. Thus personalization systems have become an important research area. This paper gives an overview of critical technologies both in the industry and academia for personalization of media and emphasizes on the state-of-art in content analysis for media recommender systems. Ling-Yu Duan, Changsheng Xu |
ICME | 1 |
| 2008 | Hierarchical movie affective content analysis based on arousal and valence featuresabstractEmotional factors directly reflect audiences' attention, evaluation and memory. Affective contents analysis not only create an index for users to access their interested movie segments, but also provide feasible entry for video highlights. Most of the work focus on emotion type detection. Besides emotion type, emotion intensity is also a significant clue for users to find their interested content. For some film genres (Horror, Action, etc), the segments with high emotion intensity have the most possibilities to be video highlights. In this paper, we propose a hierarchical structure for emotion categories and analyze emotion intensity and emotion type by using arousal and valence related features hierarchically. Firstly, High, Medium and Low are detected as emotion intensity levels by using fuzzy c-mean clustering on arousal features. Fuzzy clustering provides a mathematical model to represent vagueness, which is close to human perception. After that, valence related features are used to detect emotion types (Anger, Sad, Fear, Happy and Neutral). Considering video is continuous time series data and the occurrence of a certain emotion is affected by recent emotional history, Hidden Markov Models (HMMs) are used to capture the context information. Experimental results shows the movie segments with high emotion intensity cover over 80% of the movie highlights in Horror and Action movies and the hierarchical method outperforms the one-step method on emotion type detection. Meanwhile, it is flexible for user to pick up their favorite affective content by choosing both emotion intensity levels and emotion types. Min Xu 0001, Jesse S. Jin, Suhuai Luo, Ling-Yu Duan |
ACM Multimedia | 4 |
| 2008 | A Multimodal Scheme for Program Segmentation and Representation in Broadcast Video StreamsabstractWith the advance of digital video recording and playback systems, the request for efficiently managing recorded TV video programs is evident so that users can readily locate and browse their favorite programs. In this paper, we propose a multimodal scheme to segment and represent TV video streams. The scheme aims to recover the temporal and structural characteristics of TV programs with visual, auditory, and textual information. In terms of visual cues, we develop a novel concept named program-oriented informative images (POIM) to identify the candidate points correlated with the boundaries of individual programs. For audio cues, a multiscale Kullback-Leibler (K-L) distance is proposed to locate audio scene changes (ASC), and accordingly ASC is aligned with video scene changes to represent candidate boundaries of programs. In addition, latent semantic analysis (LSA) is adopted to calculate the textual content similarity (TCS) between shots to model the inter-program similarity and intra-program dissimilarity in terms of speech content. Finally, we fuse the multimodal features of POIM, ASC, and TCS to detect the boundaries of programs including individual commercials (spots). Towards effective program guide and attracting content browsing, we propose a multimodal representation of individual programs by using POIM images, key frames, and textual keywords in a summarization manner. Extensive experiments are carried out over an open benchmarking dataset TRECVID 2005 corpus and promising results have been achieved. Compared with the electronic program guide (EPG), our solution provides a more generic approach to determine the exact boundaries of diverse TV programs even including dramatic spots. Jinqiao Wang, Ling-Yu Duan, Qingshan Liu 0001, Hanqing Lu, Jesse S. Jin |
IEEE Trans. Multim. | 2 |
| 2008 | Audio keywords generation for sports video analysisabstractSports video has attracted a global viewership. Research effort in this area has been focused on semantic event detection in sports video to facilitate accessing and browsing. Most of the event detection methods in sports video are based on visual features. However, being a significant component of sports video, audio may also play an important role in semantic event detection. In this paper, we have borrowed the concept of the “keyword” from the text mining domain to define a set of specific audio sounds. These specific audio sounds refer to a set of game-specific sounds with strong relationships to the actions of players, referees, commentators, and audience, which are the reference points for interesting sports events. Unlike low-level features, audio keywords can be considered as a mid-level representation, able to facilitate high-level analysis from the semantic concept point of view. Audio keywords are created from low-level audio features with learning by support vector machines. With the help of video shots, the created audio keywords can be used to detect semantic events in sports video by Hidden Markov Model (HMM) learning. Experiments on creating audio keywords and, subsequently, event detection based on audio keywords have been very encouraging. Based on the experimental results, we believe that the audio keyword is an effective representation that is able to achieve satisfying results for event detection in sports video. Application in three sports types demonstrates the practicality of the proposed method. Min Xu 0001, Changsheng Xu, Ling-Yu Duan, Jesse S. Jin, Suhuai Luo |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2007 | Robust Commercial Retrieval in Video StreamsabstractTV commercial video is a kind of informative medium. To fast and robustly index and retrieve commercial videos is of interest to commercial monitor, copyright protection, and commercial management, we propose a coarse-to-fine scheme to robustly retrieve commercial videos. Different from previous work using clip or key frames-based matching, our scheme has incorporated the commercial production knowledge to search the candidate commercial positions. Color and ordinal features are extracted for locating the exact commercial positions with dynamic time warping distance. Comparison experiments were carried out over TRECVID 2006 news videos and some videos from Chinese channels. Our scheme has achieved promising simulation results. Jinqiao Wang, Ling-Yu Duan, Qingshan Liu 0001, Hanqing Lu, Jesse S. Jin |
ICME | 2 |
| 2007 | Automatic TV Logo Detection, Tracking and Removal in Broadcast Video
Jinqiao Wang, Qingshan Liu 0001, Ling-Yu Duan, Hanqing Lu, Changsheng Xu |
MMM (2) | 3 |
| 2007 | An algorithm to estimate mean vehicle speed from MPEG Skycam video
Ping Xue 0001, Ling-Yu Duan, Qi Tian 0002 |
Multim. Tools Appl. | 3 |
| 2006 | A Mid-Level Scene Change Representation Via Audiovisual AlignmentabstractScene is a series of semantic correlated video shots. An effective scene detection depends on domain knowledge more or less. Most existing approaches try to directly detect various scene changes by applying clustering or supervised learning methods to low level audiovisual features. However, robustly detecting diverse scene changes derived from complex semantic meanings is still a challenging problem. In this paper we are focused on the association of visual signal changes (e.g. cuts, fade-in, fade-out, etc.) and audio signal changes (e.g. speaker change, background music change, etc.) to propose a mid-level scene change representation, which is meant to locate candidate scene change points by characterizing temporally uncorrelated properties of audio and visual track in the case of scene change happening. By incorporating domain knowledge, enhanced features can be further extracted to complement this representation to bridge semantic gap towards scene change detection. We utilize a camera motion estimation algorithm to detect visual signal changes. Such visual change positions are selected as time-stamp points. An alignment is performed to search for candidate audio signal change positions by multi-scale Kullback-Leibler(K-L) distance computing. Both metric-based K-L distance approach and model-based HMM are applied to determine true audio signal changes. The associated visual and audio signal changes are considered as the mid-level scene change representation. This representation has been successfully applied to detect boundaries of individual commercial in TV broadcast stream with an accuracy of around 95%. Particularly the systematic alignment approach can be utilized in video summarization. Jinqiao Wang, Ling-Yu Duan, Hanqing Lu, Jesse S. Jin, Changsheng Xu |
ICASSP (2) | 2 |
| 2006 | A Robust Method for TV Logo Tracking in Video StreamsabstractMost broadcast stations rely on TV logos to claim video content ownership or visually distinguish the broadcast from the interrupting commercial block. Detecting and tracking a TV logo is of interest to TV commercial skipping applications and logo-based broadcasting surveillance (abnormal signal is accompanied by logo absence). Pixel-wise difference computing within predetermined logo regions cannot address semi-transparent TV logos well for the blending effects of a logo itself and inconstant background images. Edge-based template matching is weak for semi-transparent ones when incomplete edges appear. In this paper we present a more robust approach to detect and track TV logos in video streams on the basis of multispectral images gradient. Instead of single frame based detection, our approach makes use of the temporal correlation of multiple consecutive frames. Since it is difficult to manually delineate logos of irregular shape, an adaptive threshold is applied to the gradient image in subpixel space to extract the logo mask. TV logo tracking is finally carried out by matching the masked region with a known template. An extensive comparison experiment has shown our proposed algorithm outperforms traditional methods such as frame difference, single frame-based edge matching. Our experimental dataset comes from part of TRECVID2005 news corpus and several Chinese TV channels with challenging TV logos Jinqiao Wang, Ling-Yu Duan, Zhenglong Li 0001, Jing Liu 0001, Hanqing Lu, Jesse S. Jin |
ICME | 2 |
| 2006 | TV Commercial Classification by using Multi-Modal Textual InformationabstractIn this paper, we propose an approach for TV commercial video classification by the categories of advertised products or services (e.g. automobiles, healthcare products, etc). Since automatic speech recognition (ASR) and optical character recognition (OCR) can deliver meaningful textual information related to products or services, TV commercial video classification is formulated as the problem of text categorization. However, there exist two challenges. Firstly, the background music of TV commercials makes ASR techniques yield erroneous and deficient output transcripts. Secondly, even if ASR and OCR could work perfectly, the limited textual information from TV commercials do not suffice to train a generic and non-overfitting text categorizer. For the first issue, our approach resorts to the external resources to expand deficient ASR and OCR transcripts. The output transcripts of ASR and OCR are parsed to yield a few keywords, on which a Web searching is executed to retrieve relevant and semantically informative articles from World Wide Web (WWW). The retrieved articles are then utilized to construct textual feature vectors and perform text categorization on behalf of commercials. For the second issue, a topic-wise document corpus is constructed from the public corpora like Reuters-21578 or from the articles manually collected from WWW for the training of text categorizers. Experimental results have shown that the proposed approach alleviates the negative effects from weak ASR/OCR performance and yield a promising classification accuracy of 80.9% Yantao Zheng, Ling-Yu Duan, Qi Tian 0002, Jesse S. Jin |
ICME | 2 |
| 2006 | Segmentation, categorization, and identification of commercial clips from TV streams using multimodal analysisabstractTV advertising is ubiquitous, perseverant, and economically vital. Millions of people's living and working habits are affected by TV commercials. In this paper, we present a multimodal ("visual + audio + text") commercial video digest scheme to segment individual commercials and carry out semantic content analysis within a detected commercial segment from TV streams.Two challenging issues are addressed. Firstly, we propose a multimodal approach to robustly detect the boundaries of individual commercials. Secondly, we attempt to classify a commercial with respect to advertised products/services. For the first, the boundary detection of individual commercials is reduced to the problem of binary classification of shot boundaries via the mid-level features derived from two concepts: Image Frames Marked with Product Information (FMPI) and Audio Scene Change Indicator (ASCI). Moreover, the accurate individual boundary enables us to perform commercial identification by clip matching via a spatial-temporal signature. For the second, commercial classification is formulated as the task of text categorization by expanding sparse texts from ASR/OCR with external knowledge. Our boundary detection has achieved a good result of F1 = 93.7% on the dataset comprising 499 individual commercials from TRECVID'05 video corpus. Commercial classification has obtained a promising accuracy of 80.9% on 141 distinct ones. Based on these achievements, various applications such as an intelligent digital TV set-top box can be accomplished to enhance the TV viewer's capabilities in monitoring and managing commercials from TV streams. Ling-Yu Duan, Jinqiao Wang, Yantao Zheng, Jesse S. Jin, Hanqing Lu, Changsheng Xu |
ACM Multimedia | 1 |
| 2006 | Live sports event detection based on broadcast video and web-casting textabstractEvent detection is essential for sports video summarization, indexing and retrieval and extensive research efforts have been devoted to this area. However, the previous approaches are heavily relying on video content itself and require the whole video content for event detection. Due to the semantic gap between low-level features and high-level events, it is difficult to come up with a generic framework to achieve a high accuracy of event detection. In addition, the dynamic structures from different sports domains further complicate the analysis and impede the implementation of live event detection systems. In this paper, we present a novel approach for event detection from the live sports game using web-casting text and broadcast video. Web-casting text is a text broadcast source for sports game and can be live captured from the web. Incorporating web-casting text into sports video analysis significantly improves the event detection accuracy. Compared with previous approaches, the proposed approach is able to: (1) detect live event only based on the partial content captured from the web and TV; (2) extract detailed event semantics and detect exact event boundary, which are very difficult or impossible to be handled by previous approaches; and (3) create personalized summary related to certain event, player or team according to user's preference. We present the framework of our approach and details of text analysis, video analysis and text/video alignment. We conducted experiments on both live games and recorded games. The results are encouraging and comparable to the manually detected events. We also give scenarios to illustrate how to apply the proposed solution to professional and consumer services. Changsheng Xu, Jinjun Wang, Kong-Wah Wan, Ling-Yu Duan |
ACM Multimedia | 5 |
| 2006 | Nonparametric motion characterization for robust classification of camera motion patternsabstractMotion characterization plays a critical role in video indexing. An effective way of characterizing camera motion facilitates the video representation, indexing and retrieval tasks. This paper describes a novel nonparametric motion representation to achieve an effective and robust recognition of parts of the video in which camera is static, or panning, or tilting, or zooming, etc. This representation employs the mean shift filtering and the vector histograms to produce a compact description of a motion field. The basic idea is to perform spatio-temporal mode-seeking in the motion feature space and use the histograms-based spatial distributions of dominant motion modes to represent a motion field. Unlike most existing approaches, which focus on the estimation of a parametric motion model from a dense optical flow field (OFF) or a block matching-based motion vector field (MVF), the proposed method combines the motion representation and machine learning techniques (e.g., support vector machines) to perform camera motion analysis from the classification point of view. The main motivation lies in the impossibility of uniformly securing a proper parametric assumption in a wide range of video scenarios. The diverse camera shot sizes and frequent occurrences of bad OFF/MVF necessitates a learning mechanism, which can not only capture the domain-independent parametric constraints, but also acquire the domain-dependent knowledge to tolerate the influence of bad OFF/MVF. In order to improve performance, we can use this learning-based method to train enhanced classifiers aiming at a certain context (i.e., shot size, neighbor OFF/MVFs, and video genre). Other visual cues (e.g., dominant color) can also be incorporated for further motion analysis. Our main aim is to use a generic feature space analysis method to explore a flexible OFF/MVF representation in a nonparametric technique, which could be fed into a learning framework to robustly capture the global motion by incorporating the context information. Results on videos with various types of content (23 191 MVFs culled from MPEG-7 dataset, and 20 000 MVFs culled from broadcast tennis, soccer, and basketball videos) are reported to validate the proposed approach. Ling-Yu Duan, Jesse S. Jin, Qi Tian 0002, Changsheng Xu |
IEEE Trans. Multim. | 1 |
| 2005 | Replay Scene Classification in Soccer Video Using Web Broadcast TextabstractThe automatic extraction of sports video highlights is a typical kind of personalized media production process. Many ways have been studied from the viewpoints of low-level audio/visual processing (e. g. detection of excited commentator speech), event detection (e. g. goal detection), etc. However, the subjectivity of highlights is an unavoidable bottleneck. The replay scene is an effective clue for highlights in broad-cast sports video due to the incorporation of video production knowledge. Most related work deals with the replay detection and/or a simple composition of all detected replays to generate highlights. Different from previous work, our work considers different flavors of different people in terms of highlight content or type through replay scenes classification. The main contributions include: 1) proposing a multi-modal (visual+ textual) approach for refined replay classification; 2) employing the sources of Broadcast Web Text (BWT) to facilitate replay content analysis. An overall accuracy of 79.9% has been achieved on seven soccer matches over seven replay categories Jinhui Dai, Ling-Yu Duan, Xiaofeng Tong, Changsheng Xu, Qi Tian 0002, Hanqing Lu, Jesse S. Jin |
ICME | 2 |
| 2005 | A Mid-level Visual Concept Generation Framework for Sports AnalysisabstractThe development of mid-level concepts helps to bridge the gap between low-level feature and high-level semantics in video analysis. Most existing work combines the customized mid-level concepts and statistical models to detect particular events. Based on broadcast sports video production knowledge, we extend our previous work to present a unified framework for mid-level concept generation in this paper. A video segment is characterized via three essential aspects: camera shot size, an object appearing in a scene, and video production technology. These three aspects clearly summarize the primary concerns in terms of a generic concept generation. Within this framework, we can flexibly and clearly define meaningful mid-level concepts towards comprehensive video content analysis, such as replay classification and the detection of events (e. g. goal, shoot, attack, foul, offside, and out of bound, etc.). Xiaofeng Tong, Ling-Yu Duan, Hanqing Lu, Changsheng Xu, Qi Tian 0002, Jesse S. Jin |
ICME | 2 |
| 2005 | Periodicity Detection of Local MotionabstractPeriodicity is useful for compact representation of periodic motion and a reasonable selection of a proper temporal scale for periodic motion analysis. In this paper, we concern the periodicity detection of local motion within an interesting region and present an approach to automatically detect the motion periodicity inherent to local motion under complex condition. The task is challenging as local motion is usually buried in clutters with global motion and noises. Most exist ing methods have assumed a static camera and a labeled moving object region. We instead apply robust local motion estimation and an object localization method to extract the object motion. The object motion is characterized by the confidence based motion probability map and the motion vectors obtained by global motion compensation. The autocorrelation series of motion energy is then carried out to locate local maximum points. With the set of indices of local maximum points, we can estimate the basic periodicity through a least-square fitting. This method has been applied to swimming videos and got encouraging results. Xiaofeng Tong, Ling-Yu Duan, Changsheng Xu, Qi Tian 0002, Hanqing Lu, Jinjun Wang, Jesse S. Jin |
ICME | 2 |
| 2005 | Automatic generation of personalized music sports videoabstractIn this paper, we propose a novel automatic approach for personalized music sports video generation. Two research challenges, semantic sports video content selection and automatic video composition, are addressed. For the first challenge, we propose to use multi-modal (audio, video and text) feature analysis and alignment to detect the semantic of events in sports video. For the second challenge, we propose video-centric and music-centric music video composition schemes to automatically generate personalized music sports video based on user's preference. The experimental results and user evaluations are promising and show that our system's generated music sports video is comparable to manually generated ones. The proposed approach greatly facilitates the automatic music sports video generation for both professionals and amateurs. Jinjun Wang, Changsheng Xu, Chng Eng Siong, Ling-Yu Duan, Kong-Wah Wan, Qi Tian 0002 |
ACM Multimedia | 4 |
| 2005 | A unified framework for semantic shot classification in sports videoabstractThe extensive amount of multimedia information available necessitates content-based video indexing and retrieval methods. Since humans tend to use high-level semantic concepts when querying and browsing multimedia databases, there is an increasing need for semantic video indexing and analysis. For this purpose, we present a unified framework for semantic shot classification in sports video, which has been widely studied due to tremendous commercial potentials. Unlike most existing approaches, which focus on clustering by aggregating shots or key-frames with similar low-level features, the proposed scheme employs supervised learning to perform a top-down video shot classification. Moreover, the supervised learning procedure is constructed on the basis of effective mid-level representations instead of exhaustive low-level features. This framework consists of three main steps: 1) identify video shot classes for each sport; 2) develop a common set of motion, color, shot length-related mid-level representations; and 3) supervised learning of the given sports video shots. It is observed that for each sport we can predefine a small number of semantic shot classes, about 5-10, which covers 90%-95% of broadcast sports video. We employ nonparametric feature space analysis to map low-level features to mid-level semantic video shot attributes such as dominant object (a player) motion, camera motion patterns, and court shape, etc. Based on the fusion of those mid-level shot attributes, we classify video shots into the predefined shot classes, each of which has clear semantic meanings. With this framework we have achieved good classification accuracy of 85%-95% on the game videos of five typical ball type sports (i.e., tennis, basketball, volleyball, soccer, and table tennis) with over 5500 shots of about 8 h. With correctly classified sports video shots, further structural and temporal analysis, such as event detection, highlight extraction, video skimming, and table of content, will be greatly facilitated. Ling-Yu Duan, Min Xu 0001, Qi Tian 0002, Changsheng Xu, Jesse S. Jin |
IEEE Trans. Multim. | 1 |
| 2004 | Mean shift based video segment representation and applications to replay detectionabstractEffective and efficient representation of the low-level features of groups of frames or shots is an important yet challenging task for video analysis and retrieval. Key frame-based representation is limited by the difficulties in shot boundary detection of gradual transitions, and the variety of methods of key frame extraction. In this paper, we employ the mean shift-based mode seeking function to develop a new approach for compact representation of the video segment. The proposed video representation is motivated by recognizing that, on the global level, humans perceive images only as a combination of the few most prominent colors. We exploit the spatiotemporal mode seeking in feature space to simulate the "subjectivity" of human decisions in video segment retrieval and identification. The effectiveness of the video representation and matching scheme is shown by initial experiments on replay detection in broadcast sports videos. Ling-Yu Duan, Min Xu 0001, Qi Tian 0002, Changsheng Xu |
ICASSP (5) | 1 |
| 2004 | Mean shift based nonparametric motion characterization
Ling-Yu Duan, Min Xu 0001, Qi Tian 0002, Changsheng Xu |
ICIP | 1 |
| 2004 | Nonparametric motion model with applications to camera motion pattern classificationabstractMotion information is a powerful cue for visual perception. In the context of video indexing and retrieval, motion content serves as a useful source for compact video representation. There has been a lot of literature about parametric motion models. However, it is hard to secure a proper parametric assumption in a wide range of video scenarios. Diverse camera shots and frequent occurrences of bad optical flow estimation motivate us to develop nonparametric motion models. In this paper, we employ the mean shift procedure to propose a novel nonparametric motion representation. With this compact representation, various motion characterization tasks can be achieved by machine learning. Such a learning mechanism can not only capture the domain-independent parametric constraints, but also acquire the domain-dependent knowledge to tolerate the influence of bad dense optical flow vectors or block-based MPEG motion vector fields (MVF). The proposed nonparametric motion model has been applied to camera motion pattern classification on 23191 MVF extracted from MPEG-7 dataset. Ling-Yu Duan, Min Xu 0001, Qi Tian 0002, Changsheng Xu |
ACM Multimedia | 1 |
| 2004 | Nonparametric motion modelabstractMotion information is a powerful cue for visual perception. In the context of video indexing and retrieval, motion content serves as a useful source for compact video representation. There has been a lot of literature about parametric motion models. However, it is hard to secure a proper parametric assumption in a wide range of video scenarios. Diverse camera shots and frequent occurrences of improper optical flow estimation or block matching motivate us to develop nonparametric motion models. In this demonstration, we present a novel nonparametric motion model. The unique features mainly include: 1) Instead of computationally expensive and vulnerable parametric regression our proposed model bases the motion characterization on the classification of motion patterns; 2) we employ machine learning to capture the knowledge of recognizing camera motion patterns from bad motion vector fields (MVF); and 3) with the mean shift filtering our proposed motion representation elegantly incorporates the spatial-range information for noise removal and discontinuity preserving smoothing of MVF. Promising results have been achieved on two tasks: 1) camera motion pattern recognition on 23191 MVFs and 2) recognition of the intensity of motion activity on 622 video segments culled from the MPEG-7 dataset. Ling-Yu Duan, Min Xu 0001, Qi Tian 0002, Changsheng Xu |
ACM Multimedia | 1 |
| 2004 | Fast and robust video clip search using index structureabstractContent based retrieval of similar multimedia objects (e.g. images, text, and videos) is an important research issue in the field of multimedia database. In this demo, we present a fast and robust video clip searching system. This system consists of two major modules, namely, robust video representation and fast searching. Different from traditional key frame-based histogram methods, we employ the cumulative histogram to represent the ordinal features and color range features for a video segment. This representation provides a spatio-temporal description of the whole segment. Our experiment has shown it is effective for capturing the patterns of short video clips such as commercial, program lead in/out, flying logo in sports video, etc. In order to improve the performance in searching large video database, we introduce the index structure to deal with video search from the viewpoint of query processing (e.g. K-NN query, Range query, etc.) in high-dimensional spaces. Different query processing support different search tasks. In this demo, we employ the mrkd-tree index structure and the proposed video representation to fulfill fast and robust search of short video clips (i.e. news video lead-in/out, replay logo, commercial) in large video collections with the total length of 15 hours. Ling-Yu Duan, Junsong Yuan 0001, Qi Tian 0002, Changsheng Xu |
ACM Multimedia | 1 |
| 2004 | Audio keyword generation for sports video analysisabstractSemantic sports video analysis has attracted many research interests and audio cues have been shown to play an important role in semantics inference. To facilitate event detection using audio information, we have introduced the concept of audio keyword (e.g. excited/plain commentator speech, excited/plain audience sound, etc.) to describe the game-specific sound associated with an event. In our previous work, we have designed a hierarchical Support Vector Machine (SVM) classifier for audio keyword identification. However, there are two inherent weaknesses: 1) a frame-based SVM classifier does not incorporate any contextual information; 2) a robust recognizer relies on large amounts of training data in the case of different sports games videos. In this demo, we present a flexible Hidden Markov Model (HMM)-based audio keyword generation system. This is motivated by the successful story of applying HMM in speech recognition. Unlike the frame-based SVM classification followed by a major voting, our HMM-based system treats an audio keyword as a continuous time series data and employs hidden states transition to capture contexts. Moreover, our system introduces an adaptation mechanism to tune the initial HMM models (obtained from available training data) to improve performance by a small number of data from a new sports game video. Promising results has been demonstrated on the tennis, soccer and basketball videos with the total length of 2 hours. Min Xu 0001, Ling-Yu Duan, Liang-Tien Chia, Changsheng Xu |
ACM Multimedia | 2 |
| 2003 | A fusion scheme of visual and auditory modalities for event detection in sports videoabstractWe propose an effective fusion scheme of visual and auditory modalities to detect events in sports video. The proposed scheme is built upon semantic shot classification, where we classify video shots into several major or interesting classes, each of which has clear semantic meanings. Among major shot classes we perform classification of the different auditory signal segments (i.e. silence, hitting ball, applause, commentator speech) with the goal of detecting events with strong semantic meaning. For instance, for tennis video, we have identified five interesting events: serve, reserve, ace, return, and score. Since we have developed a unified framework for semantic shot classification in sports videos and a set of audio mid-level representation with supervised learning methods, the proposed fusion scheme can be easily adapted to a new sports game. We are extending this fusion scheme to three additional typical sports videos: basketball, volleyball and soccer. Correctly detected sports video events will greatly facilitate further structural and temporal analysis, such as sports video skimming, table of content, etc. Min Xu 0001, Ling-Yu Duan, Changsheng Xu, Qi Tian 0002 |
ICASSP (3) | 2 |
| 2003 | Robust moving video object segmentation in the MPEG compressed domainabstractIn this paper, we proposed a robust moving video object segmentation algorithm using features in the MPEG compressed domain. We first cluster the motion vectors and produce a motion mask. Then, a difference mask at 8 x 8 block size is extracted from the DC image by background subtraction method. Finally, the motion mask and the difference mask are combined conditionally to generate the final object mask based on a set of rules that are application specified and is obtained with learning or heuristic based methods. The experimental results show that this object segmentation scheme is more robust than those using DC image or motion vectors only. Xiao-Dong Yu, Ling-Yu Duan, Qi Tian 0002 |
ICIP (3) | 2 |
| 2003 | A fusion scheme of visual and auditory modalities for event detection in sports videoabstractIn this paper, we propose an effective fusion scheme of visual and auditory modalities to detect events in sports video. The proposed scheme is built upon semantic shot classification, where we classify video shots into several major or interesting classes, each of which has clear semantic meanings. Among major shot classes we perform classification of the different auditory signal segments (i.e. silence, hitting ball, applause, commentator speech) with the goal of detecting events with strong semantic meaning. For instance, for tennis video, we have identified five interesting events: serve, reserve, ace, return, and score. Since we have developed a unified framework for semantic shot classification in sports videos and a set of audio mid-level representation with supervised learning methods, the proposed fusion scheme can be easily adapted to a new sports game. We are extending this fusion scheme to three additional typical sports videos: basketball, volleyball and soccer. Correctly detected sports video events will greatly facilitate further structural and temporal analysis, such as sports video skimming, table of content, etc. Min Xu 0001, Ling-Yu Duan, Changsheng Xu, Qi Tian 0002 |
ICME | 2 |
| 2003 | A mid-level representation framework for semantic sports video analysisabstractSports video has been widely studied due to its tremendous commercial potentials. Despite encouraging results from various specific sports games, it is almost impossible to extend a system for a new sports game because they usually employ different sets of low-level features appropriate for the specific games and closely coupled with the use of game specific rules to detect events or highlights. There is a lack of internal representation and structure to be generic and applicable for many different sports. In this paper, we present a generic mid-level representation framework for semantic sports video analysis. The mid-level representation layer is introduced between the low-level audiovisual processing and high-level semantic analysis. It allows us to separate sports specific knowledge and rules from the low-level and mid-level feature extraction. This makes sports video analysis more efficient, effective, and less ad-hoc for various types of sports. To achieve robustness of the low-level feature analysis, a non-parametric clustering, mean shift procedure, has been successfully applied to both color and motion analysis. The proposed framework has been tested for five field-ball type sports covering duration of about 8 hours. Experiments have shown its robust performance in semantic analysis and event detection. We believe that the proposed mid-level representation framework can be used for event detection, highlight extraction, summarization and personalization of many types of sports video. Categories and Subject Descriptors H.3.1 [Content Analysis and Indexing]: abstracting methods, indexing methods. Ling-Yu Duan, Min Xu 0001, Tat-Seng Chua, Qi Tian 0002, Changsheng Xu |
ACM Multimedia | 1 |
| 2003 | Nonparametric color characterization using mean shiftabstractColor is very useful in locating and recognizing objects that occur in artificial environments. The color histogram has shown its efficiency and advantages as a general tool for various applications, such as content-based image retrieval and video browsing, object indexing and location, and video segmentation. However, due to the lack of any spatial and context information, the histogram is not robust and effective for color characterization (e.g. dominant color) in large video databases. In this paper, we propose a nonparametric color characterization model using mean shift procedure, with an emphasis on spatio-temporal consistency. Experimental results suggest that the color characterization model is much more effective for video indexing and browsing, particularly in the domain of structured video (e.g. sports video). Ling-Yu Duan, Min Xu 0001, Qi Tian 0002, Changsheng Xu |
ACM Multimedia | 1 |
| 2002 | Clear face analysis from MPEG compressed videoabstractIn this demonstration, we present a system to analyze the clear degree of faces present in MPEG compressed video of Head-and-Shoulders style. The proposed system consists of three hierarchical modules: low-level features extraction, robust face tracking, and clear faces selection. We have integrated the core algorithm into an Automated Transaction Service (ATS) surveillance system. The Incremental Focus of Attention (IFA) architecture is taken to combine pixel domain processing with compressed domain processing --- thus, implemented system exhibits computational efficiency and tolerance to very cluttered scenes. The proposed system has successively detected segments with clear frontal faces from more than 20 Automated Teller Machine (ATM) testing clips in MPEG format, each of which consists of 1~3 transactions. In addition, the proposed scheme implies some potential video mining applications, such as automatic checking to verify entry authorization, retrieval of suspicious activities in prerecorded video surveillance sequences. Ling-Yu Duan, Qi Tian 0002 |
ACM Multimedia | 1 |
| 2002 | A unified framework for semantic shot classification in sports videosabstractIn this demonstration, we present a unified framework for semantic shot classification in sports videos. Unlike previous approaches, which focus on clustering by aggregating shots with similar low-level features, the proposed scheme makes use of domain knowledge of specific sport to perform a top-down video shot classification, including identification of video shots classes for each sport, and supervised learning and classification of given sports video with low-level and middle-level features extracted from the sports video. It's observed that for each sport we can predefine a small number of semantic shot classes, 5--10, which cover 90 to 95 % of sports broadcasting video. With supervised learning method, we can map the low-level features to middle-level semantic video shot attributes such as dominant object motion (a player), camera motion patterns, and court shape, etc. On the basis of the appropriate fusion of those middle-level shot attributes, we classify video shots into the predefined video shot classes, each of which has a clear semantic meaning. The proposed method has been tested over 3 types of sports videos: tennis, basketball, and soccer. Good classification results ranging from 80~95% have been achieved. The proposed framework provides a generic solution for sports video semantic shot classification, which can be adapted to a new sport type easily. With correctly classified sports video shots further structural and temporal analysis will be greatly facilitated. Ling-Yu Duan, Min Xu 0001, Xiao-Dong Yu, Qi Tian 0002 |
ACM Multimedia | 1 |