Jian Cheng 0001

dblp:14/6145-1 · DBLP profile ↗
← Back
252ranked-venue papers
20as first author
88since 2021 · last 2026
0000-0003-1289-2758ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 149 · 13 first-author · 59 since 2021Graphics, computer vision, multimedia, augmented reality and games · 132 · 7 first-author · 34 since 2021Systems, architecture and hardware · 17 · 14 since 2021Databases, data management, data science and information retrieval · 17 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 8 · 6 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
YearPublicationVenuePosition
2026 Deep (Predictive) Discounted Counterfactual Regret Minimization
abstract
Counterfactual regret minimization (CFR) is a family of algorithms for effectively solving imperfect-information games. To enhance CFR's applicability in large games, researchers use neural networks to approximate its behavior. However, existing methods are mainly based on vanilla CFR and struggle to effectively integrate more advanced CFR variants. In this work, we propose an efficient model-free neural CFR algorithm, overcoming the limitations of existing methods in approximating advanced CFR variants. At each iteration, it collects variance-reduced sampled advantages based on a value network, fits cumulative advantages by bootstrapping, and applies discounting and clipping operations to simulate the update mechanisms of advanced CFR variants. Experimental results show that, compared with model-free neural algorithms, it exhibits faster convergence in typical imperfect-information games and demonstrates stronger adversarial performance in a large poker game.
Hang Xu 0006, Kai Li 0022, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
AAAI6
2026 Boosting the Performance of Tree-Based Speculative Decoding of LLMs on FPGAs
abstract
As an efficient alternative to autoregressive decoding, tree-based speculative decoding (SD) has been widely adopted to accelerate LLM inference on GPUs. However, due to the notable disparity in compute power and memory bandwidth, we observe that a specific target-draft model pair with a proper decoding configuration, despite demonstrating significant performance gains on GPUs, often fails to maintain its efficacy on FPGAs, and may even underperform the standard autoregressive decoding approachIn this paper, we propose an analytical framework to revive the performance of tree-based speculative decoding on FPGAs. We introduce effective performance, a roofline-based metric designed to: 1) assess whether a specific target-draft model pair can benefit from SD for the given FPGA platform, and 2) determine the optimal decoding configuration to achieve peak performance when SD is applicable. We also propose a prior-score-based search strategy to identify the optimal tree structure for a preset number of nodes, further enhancing the performance. We evaluate our method on AMD FPGA platforms using two state-of-the-art SD algorithms: LongSpec and EAGLE-3. Our approach demonstrates a speedup of 2.54-3.89× over autoregressive decoding.
Tielong Liu, Gang Li 0015, Zitao Mo, Minnan Pei, Jian Cheng 0001
DATE6
2026 APEX: Integer-only Non-linear Function Approximation for Efficient Cross-Modal Inference
Peihuan Ni, Zitao Mo, Tielong Liu, Hongli Wen, Minnan Pei, Junwen Si, Weifan Guan, Peisong Wang 0001, Qinghao Hu 0001, Gang Li 0015, Jian Cheng 0001
DATE12
2026 KASH: 1-Bit Key-Value Cache Quantizaiton via Asymmetric Hashing
Qinghao Hu 0001, Zhihui Wei, Jian Cheng 0001
ICPR (14)4
2026 A Universal Self-Attention Enhancement for Bridging Low-bit Quantization and Vision Transformers
abstract
Low-bit quantization of Vision Transformers presents significant challenges due to the intrinsic properties of their Multi-head Self-attention modules. In this work, we investigate the quantization issues specific to this mechanism and identify two critical challenges. First, quantization amplifies discrepancies among attention heads, thereby impairing the model’s capability to focus on the most informative regions. Second, high-magnitude values in the attention maps, which encode essential relational information, exhibit an extremely sharp distribution that renders them especially prone to substantial information loss during quantization. To address these challenges, we propose Quantized-aware Multi-head Self-attention (Q-MHSA), a universal self-attention enhancement module that integrates two lightweight components within Multi-head Self-attention (MHSA). The Cross-head Concordance Module (CCM) enforces adaptive consistency across attention heads, while the Learnable Smoothness Controller (LSC) replaces the fixed normalization factor with an adaptive mechanism that selectively smooths the distribution of high-magnitude attention values while disregarding less informative low values. Designed for seamless integration with any Vision Transformer architecture and quantization-aware training method, Q-MHSA incurs minimal overhead while consistently improving model accuracy. For instance, under the VVTQ quantization framework applied to a 4-bit Swin-S model, the incorporation of Q-MHSA yields a top-1 accuracy of 83.5%, representing a 0.9% improvement over the baseline, while incurring a marginal overhead of 0.01% in both parameters and FLOPs.
Jiahe Qian, Peisong Wang 0001, Zhengyang Zhuge, Qinghao Hu 0001, Jian Cheng 0001
WACV5
2026 Towards efficient and accurate spiking neural networks via adaptive bit allocation
abstract
Multi-bit spiking neural networks (SNNs) have recently become a heated research spot, pursuing energy-efficient and high-accurate AI. However, with more bits involved, the associated memory and computation demands escalate to the point where the performance improvements become disproportionate. Based on the insight that different layers demonstrate different importance and extra bits could be wasted and interfering, this paper presents an adaptive bit allocation strategy for direct-trained SNNs, achieving fine-grained layer-wise allocation of memory and computation resources. Thus, SNN's efficiency and accuracy can be improved. Specifically, we parametrize the temporal lengths and the bit widths of weights and spikes, and make them learnable and controllable through gradients. To address the challenges caused by changeable bit widths and temporal lengths, we propose the refined spiking neuron, which can handle different temporal lengths, enable the derivation of gradients for temporal lengths, and suit spike quantization better. In addition, we theoretically formulate the step-size mismatch problem of learnable bit widths, which may incur severe quantization errors to SNN, and accordingly propose the step-size renewal mechanism to alleviate this issue. Experiments on various datasets, including the static CIFAR and ImageNet datasets and the dynamic CIFAR-DVS and DVS-GESTURE datasets, demonstrate that our methods can reduce the overall memory and computation cost while achieving higher accuracy. Particularly, our SEWResNet-34 can achieve a 2.69 % accuracy gain and 4.16 × lower bit budgets over the advanced baseline work on ImageNet. This work is open-sourced at this link.
Xingting Yao, Qinghao Hu 0001, Tielong Liu, Gang Li 0015, Peisong Wang 0001, Jian Cheng 0001
Neural Networks7
2026 Dual Assignment of labels for end-to-end fully convolutional object detection
Qiang Chen 0007, Qinghao Hu 0001, Jian Cheng 0001
Pattern Recognit.4
2026 MATA: A Memory-Efficient Attention Accelerator for LLMs Exploiting Look-Back KV Cache Pruning
abstract
Transformer-based Large Language Models (LLMs) have sparked a new wave of AI applications. However, their large computational complexity and memory footprint pose significant challenges for real-world deployment. Although dedicated transformer accelerators have been widely explored, we observe that they are unefficient for decoder-only LLMs that feature autoregressive computations with KV Cache. Our in-depth analysis reveals that DRAM accesses induced by the KV Cache dominate the overall attention process. To address this issue, we propose aMemory-efficientATtentionAccelerator (MATA) for LLMs through algorithm and hardware co-design. Specifically,at the algorithm level, to mitigate the overhead caused by the linear increase of KV Cache, we propose a post-training Look-Back pruning method. It dynamically discards unimportant tokens through a comprehensive scoring scheme, thereby restricting KV Cache to a constant volume.At the hardware level, to identify important tokens with low latency, we design a Slice Top-K (STK) engine that can complete top-k-based sorting withO(N) time complexity. Moreover, we present the Adaptive Dataflow, which adaptively performs different inference phases of LLMs, thus significantly enhancing the PE array utilization. On average, our MATA can achieve speedups of 3.56×, 2.23× and 2.04×, 1.56× energy savings over two state-of-the-art transformer accelerators SpAtten and FACT, respectively.
Gang Li 0015, Tielong Liu, Zitao Mo, Xiaoyao Liang, Jian Cheng 0001
IEEE Trans. Computers6
2025 An Open-Ended Learning Framework for Opponent Modeling
abstract
Opponent Modeling (OM) aims to enhance decision-making by modeling other agents in multi-agent environments. Existing works typically learn opponent models against a pre-designated fixed set of opponents during training. However, this will cause poor generalization when facing unknown opponents during testing, as previously unseen opponents can exhibit out-of-distribution (OOD) behaviors that the learned opponent models cannot handle. To tackle this problem, we introduce a novel Open-Ended Opponent Modeling (OEOM) framework, which continuously generates opponents with diverse strengths and styles to reduce the possibility of OOD situations occurring during testing. Founded on population-based training and information-theoretic trajectory space diversity regularization, OEOM generates a dynamic set of opponents. This set is then fed to any OM approaches to train a potentially generalizable opponent model. Upon this, we further propose a simple yet effective OM approach that naturally fits within the OEOM framework. This approach is based on in-context reinforcement learning and learns a Transformer that dynamically recognizes and responds to opponents based on their trajectories. Extensive experiments in cooperative, competitive, and mixed environments demonstrate that OEOM is an approach-agnostic framework that improves generalizability compared to training against a fixed set of opponents, regardless of OM approaches or testing opponent settings. The results also indicate that our proposed approach generally outperforms existing OM baselines.
Yuheng Jing, Kai Li 0022, Bingyun Liu, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
AAAI7
2025 EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
abstract
Mixture-of-Experts (MoE) has demonstrated promising potential in scaling LLMs.However, it is hindered by two critical challenges: (1) substantial GPU memory consumption to load all experts; (2) low activated parameters cannot be equivalently translated into inference acceleration effects.In this work, we propose EAC-MoE, an Expert-Selection Aware Compressor for MoE-LLMs, which deeply aligns with the characteristics of MoE from the perspectives of quantization and pruning, and introduces two modules to address these two challenges respectively: (1) The expert selection bias caused by low-bit quantization is a major factor contributing to the performance degradation in MoE-LLMs.Based on this, we propose Quantization with Expert-Selection Calibration (QESC), which mitigates the expert selection bias by calibrating the routers within the MoE;(2) There are always certain experts that are not crucial for the corresponding tasks, yet causing inference latency.Therefore, we propose Pruning based on Expert-Selection Frequency (PESF), which significantly improves inference speed by pruning less frequently used experts for current task.Extensive experiments demonstrate that our approach significantly reduces memory usage and improves inference speed with minimal performance degradation.
Yuanteng Chen, Yuantian Shao, Peisong Wang 0001, Jian Cheng 0001
ACL (1)4
2025 Exploring Contextual Attribute Density in Referring Expression Counting
abstract
Referring expression counting (REC) algorithms are for more flexible and interactive counting ability across varied fine-grained text expressions. However, the requirement for fine-grained attribute understanding poses challenges for prior arts, as they struggle to accurately align attribute information with correct visual patterns. Given the proven importance of “visual density”, it is presumed that the limitations of current REC approaches stem from an under-exploration of “contextual attribute density” (CAD). In the scope of REC, we define CAD as the measure of the information intensity of one certain fine-grained attribute in visual regions. To model the CAD, we propose a U- shape CAD estimator in which referring expression and multi-scale visual features from GroundingDINO can interact with each other With additional density supervision, we can effectively encode CAD, which is subsequently decoded via a novel attention procedure with CAD-refined queries. Integrating all these contributions, our framework significantly outperforms state-of-the-art REC methods, achieves 30% error reduction in counting metrics and a 10% improvement in localization accuracy. The surprising results shed light on the significance of contextual attribute density for REC. Code will be at github.com/Xu3XiWang/CAD-GD.
Zhicheng Wang 0002, Jian Cheng 0001, Liwen Xiao, Zhiguo Cao 0001
CVPR4
2025 SBQ: Exploiting Significant Bits for Efficient and Accurate Post-Training DNN Quantization
abstract
Post-Training Quantization is an effective technique for deep neural network acceleration. However, as the bit-width decreases to 4 bits and below, PTQ faces significant challenges in preserving accuracy, especially for attention-based models like LLMs. The main issue lies in considerable clipping and rounding errors induced by the limited number of quantization levels and narrow data range in conventional low-precision quantization. In this paper, we present an efficient and accurate PTQ method that targets 4 bits and below through algorithm and architecture co-design. Our key idea is to dynamically extract a small portion of significant bit terms from high-precision operands to perform low-precision multiplications under the given computational budget. Specifically, we propose Significant-Bit Quantization (SBQ). It exploits a product-aware method to dynamically identify significant terms and an error-compensated computation scheme to minimize compute errors. We present a dedicated inference engine to unleash the power of SBQ. Experiments on CNNs, ViTs, and LLMs reveal that SBQ consistently outperforms prior PTQ methods under 2~4-bit quantization. We also compare the proposed inference engine with state-of-the-art bit-operation-based quantization architectures TQ and Sibia. Results show that SBQ can achieve the highest area and energy efficiency.
Jiayao Ling, Gang Li 0015, Qinghao Hu 0001, Xiaolong Lin, Jian Cheng 0001, Xiaoyao Liang
DATE6
2025 Light-DiT: An Importance-Aware Dynamic Compression Framework for Diffusion Transformers
Gang Li 0015, Xuan Zhang 0001, Jiayao Ling, Xiaolong Lin, Zhuoran Song, Jian Cheng 0001, Xiaoyao Liang
Euro-Par (2)7
2025 Maximum Entropy Reinforcement Learning with Diffusion Policy
abstract
The Soft Actor-Critic (SAC) algorithm with a Gaussian policy has become a mainstream implementation for realizing the Maximum Entropy Reinforcement Learning (MaxEnt RL) objective, which incorporates entropy maximization to encourage exploration and enhance policy robustness. While the Gaussian policy performs well on simpler tasks, its exploration capacity and potential performance in complex multi-goal RL environments are limited by its inherent unimodality. In this paper, we employ the diffusion model, a powerful generative model capable of capturing complex multimodal distributions, as the policy representation to fulfill the MaxEnt RL objective, developing a method named MaxEnt RL with Diffusion Policy (MaxEntDP). Our method enables efficient exploration and brings the policy closer to the optimal MaxEnt policy. Experimental results on Mujoco benchmarks show that MaxEntDP outperforms the Gaussian policy and other generative models within the MaxEnt RL framework, and performs comparably to other state-of-the-art diffusion-based online RL algorithms. Our code is available at https://github.com/diffusionyes/MaxEntDP.
Xiaoyi Dong, Jian Cheng 0001, Xi Sheryl Zhang
ICML2
2025 Goal-Oriented Skill Abstraction for Offline Multi-Task Reinforcement Learning
abstract
Offline multi-task reinforcement learning aims to learn a unified policy capable of solving multiple tasks using only pre-collected task-mixed datasets, without requiring any online interaction with the environment. However, it faces significant challenges in effectively sharing knowledge across tasks. Inspired by the efficient knowledge abstraction observed in human learning, we propose Goal-Oriented Skill Abstraction (GO-Skill), a novel approach designed to extract and utilize reusable skills to enhance knowledge transfer and task performance. Our approach uncovers reusable skills through a goal-oriented skill extraction process and leverages vector quantization to construct a discrete skill library. To mitigate class imbalances between broadly applicable and task-specific skills, we introduce a skill enhancement phase to refine the extracted skills. Furthermore, we integrate these skills using hierarchical policy learning, enabling the construction of a high-level policy that dynamically orchestrates discrete skills to accomplish specific tasks. Extensive experiments on diverse robotic manipulation tasks within the MetaWorld benchmark demonstrate the effectiveness and versatility of GO-Skill.
Jinmin He, Kai Li 0022, Yifan Zang 0001, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
ICML7
2025 Offline Opponent Modeling with Truncated Q-driven Instant Policy Refinement
abstract
Offline Opponent Modeling (OOM) aims to learn an adaptive autonomous agent policy that dynamically adapts to opponents using an offline dataset from multi-agent games. Previous work assumes that the dataset is optimal. However, this assumption is difficult to satisfy in the real world. When the dataset is suboptimal, existing approaches struggle to work. To tackle this issue, we propose a simple and general algorithmic improvement framework, Truncated Q-driven Instant Policy Refinement (TIPR), to handle the suboptimality of OOM algorithms induced by datasets. The TIPR framework is plug-and-play in nature. Compared to original OOM algorithms, it requires only two extra steps: (1) Learn a horizon-truncated in-context action-value function, namely Truncated Q, using the offline dataset. The Truncated Q estimates the expected return within a fixed, truncated horizon and is conditioned on opponent information. (2) Use the learned Truncated Q to instantly decide whether to perform policy refinement and to generate policy after refinement during testing. Theoretically, we analyze the rationale of Truncated Q from the perspective of No Maximization Bias probability. Empirically, we conduct extensive comparison and ablation experiments in four representative competitive environments. TIPR effectively improves various OOM algorithms pretrained with suboptimal datasets.
Yuheng Jing, Kai Li 0022, Bingyun Liu, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
ICML8
2025 LISLLM: Long Context Inference of Large Language Models with Short KV Cache
abstract
Large Language Models (LLMs) have demonstrated remarkable performance across various language tasks. However, in the process of generating long texts, the linearly increasing keyvalue (KV) cache imposes a large volume of memory footprint on HBM, resulting in significant time and energy consumption during inference. By retaining only the initial and recent tokens, existing methods ignore the intermediate tokens and the effect of the distance between current tokens and those in the KV cache, which causes significant accuracy degradation and limits the potential for KV cache compression. To overcome the above problems, we propose LISLLM, a KV cache compression mechanism that takes both the difference in importance and the distance between tokens into consideration for efficient inference of LLMs in long-context settings. Compared to the state-of-theart method, our method achieves compression ratios of up to$1.91 \times$and speedup of$1.20 \times$with better accuracy.
Tielong Liu, Gang Li 0015, Zitao Mo, Xingting Yao, Jian Cheng 0001
ICPADS6
2025 Pattern Extraction Learning for Cooperative Multi-Agent Reinforcement Learning
abstract
Multi-agent reinforcement learning (MARL) has demonstrated its superiority in addressing complex decision-making tasks involving multiple agents. However, the intricate and dynamic interactions among agents make this problem exceptionally challenging. Existing MARL methods often simplify the problem by implicitly decomposing shared rewards into individual utilities, neglecting the underlying interconnections between relevant entities. To overcome these limitations, we propose a novel framework of Multi-Agent Pattern Extraction (MAPE), which captures cooperation patterns from both agent-level and global perspectives to enhance decision-making and collaboration efficiency. Specifically, MAPE introduces two key modules: the Agent Pattern Extractor (APE) and the Global Pattern Extractor (GPE) that focus on specific interactions between the entities of interests from individual and global perspectives, respectively. The APE module focuses on computing each agent’s attention to other entities across different interaction patterns, providing this information to the GPE module. The GPE module then integrates the agent-specific pattern information and state features to identify the overall interaction pattern of the multi-agent system. By filtering out irrelevant interactions between unrelated entities and highlighting meaningful relationships, MAPE fosters more focused cooperation and facilitates more efficient learning. Extensive experiments on the StarCraft II micromanagement benchmark showcase the effectiveness of MAPE in improving efficiency in complex multi-agent environments.
Yifan Zang 0001, Jinmin He, Kai Li 0022, Junliang Xing, Jian Cheng 0001
IJCNN5
2025 GCC: A 3DGS Inference Architecture with Gaussian-Wise and Cross-Stage Conditional Processing
abstract
3D Gaussian Splatting (3DGS) has emerged as a leading neural rendering technique for high-fidelity view synthesis, prompting the development of dedicated 3DGS accelerators for resource-constrained platforms.The conventional decoupled preprocessing-rendering dataflow in existing accelerators has two major limitations: 1) a
Minnan Pei, Gang Li 0015, Junwen Si, Zitao Mo, Peisong Wang 0001, Zhuoran Song, Xiaoyao Liang, Jian Cheng 0001
MICRO9
2025 DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
abstract
Quantization plays a crucial role in accelerating the inference of large-scale models, and rotational matrices have been shown to effectively improve quantization performance by smoothing outliers. However, end-to-end fine-tuning of rotational optimization algorithms incurs high computational costs and is prone to overfitting. To address this challenge, we propose an efficient distribution-aware rotational calibration method, DartQuant, which reduces the complexity of rotational optimization by constraining the distribution of the activations after rotation. This approach also effectively reduces reliance on task-specific losses, thereby mitigating the risk of overfitting. Additionally, we introduce the QR-Orth optimization scheme, which replaces expensive alternating optimization with a more efficient solution. In a variety of model quantization experiments, DartQuant demonstrates superior performance. Compared to existing methods, it achieves 47$\times$ acceleration and 10$\times$ memory savings for rotational optimization on a 70B model. Furthermore, it is the first to successfully complete rotational calibration for a 70B model on a single 3090 GPU, making quantization of large language models feasible in resource-constrained environments.
Yuantian Shao, Yuanteng Chen, Peisong Wang 0001, Jianlin Yu, Yiwu Yao, Zhihui Wei, Jian Cheng 0001
NeurIPS8
2025 Bi-Level Knowledge Transfer for Multi-Task Multi-Agent Reinforcement Learning
abstract
Multi-Agent Reinforcement Learning (MARL) has achieved remarkable success in various real-world scenarios, but its high cost of online training makes it impractical to learn each task from scratch. To enable effective policy reuse, we consider the problem of zero-shot generalization from offline data across multiple tasks. While prior work focuses on transferring individual skills of agents, we argue that the effective policy transfer across tasks should also capture the team-level coordination knowledge. In this paper, we propose Bi-Level Knowledge Transfer (BiKT) for Multi-Task MARL, which performs knowledge transfer at both the individual and team levels. At the individual level, we extract transferable individual skill embeddings from offline MARL trajectories. At the team level, we define tactics as coordinated patterns of skill combinations and capture them by leveraging the learned skill embeddings. We map skill combinations into compact tactic embeddings and then construct a tactic codebook. To incorporate both skills and tactics into decision-making, we design a bi-level decision transformer that infers them in sequence. Our BiKT leverages both the generalizability of individual skills and the diversity of tactics, enabling the learned policy to perform effectively across multiple tasks. Extensive experiments on SMAC and MPE benchmarks demonstrate that BiKT achieves strong generalization to previously unseen tasks.
Jinmin He, Yifan Zhang 0001, Yifan Zang 0001, Jian Cheng 0001
NeurIPS6
2025 Guest Editorial: Multi-view representation learning for computer vision
abstract
Object recognition and scene analysis in single-view images may face difficulties such as occlusion and incomplete information, while multi-view learning can address this limitation. When an object or scene is observed from multiple views, information on target objects can be significantly enriched to improve the performance of computer vision tasks. For this reason, multi-view has become one of the important forms of data representation, which leads to the emerging of new research topics on complete or in-complete multi-view learning. Multi-view learning enables the use of multi-source information, nevertheless, the heterogeneous characteristics of data make it difficult to reliably associate information from different views, especially in a complex environment. It remains challenging for tasks to make effective use of the consistent and complementary information between different complete views and to enhance the completeness of potential representation. A wide variety of research is being conducted to explore and discover possible challenges and opportunities to exploit multi-view representation learning for computer vision. The purpose of this Special Issue is to collect high-quality articles on the recent development and trend of multi-view representation learning in computer vision, publish new ideas, theories, solutions and insights on this topic, and showcase their applications. In this Special Issue, we have received 36 papers, all of which underwent peer review. Of the 36 originally submitted papers, 10 have been accepted, which cover a variety of fields, such as person re-identification, gait recognition, 3D object recognition, and behaviour recognition. These accepted papers are mainly divided into three categories. The first category covers the incomplete multi-view data learning theoretics and methods. The papers in this category are of He et al., Kun et al., Fan et al. and Wang et al. The last two categories are both multi-view applications. One of which is 3D-related applications. The papers in this category are of Qi et al. and Sun et al. The other category is about 2D recognition. The papers in this category are of Zhang et al., Huang et al., Zheng et al. and Zhang et al. A brief presentation of each of the paper follows. He et al. present an innovative multi-view subspace clustering method with incomplete graph information. Specifically, they separate one shared and multiple specific graphs from multiple raw graph data, and exploit the mask fusion strategy and block diagonal regulariser to obtain the inherent category information. The clustering results on six real-world datasets show that the method outperforms a series of classic incomplete multi-view clustering methods. Kun et al. propose a new method for low-rank-based multi-view subspace clustering based on low-rank correlation analysis. To overcome the limitations of unreliable low-rank structure and imprecise graphs caused by multi-view noise and outliers, they introduce the canonical correlation analysis strategy and a dual regularisation term to characterise the connections between different views adaptively. Experimental results reveal the method's superiority over compared state-of-the-art (SOTA) methods in accuracy, normalised mutual information, and F-score evaluation metrics. Fan et al. address the challenge of partial mapping between the views in multi-view clustering, and propose a self-inferring incomplete multi-view clustering algorithm to explore the information hidden in the local geometric structure and recover missing instances through mining the information hidden in existing instances. Experimental results show that the method can improve the clustering performance compared with the SOTA methods. Wang et al. propose a semi-paired semi-supervised deep hashing to solve the large-scale multimedia retrieval task. The method is an end-to-end deep neural network model with high-order affinity. To maintain the consistency within the modalities, they introduce a common representation that combines with the labelled information to associate different modalities. Experimental results demonstrate the superior performance of proposed method. Qi et al. propose a double-weighting convolution neural network based on the L2-S grouping mechanism for multi-view 3D object recognition. The goal of the proposed L2-S grouping mechanism is to calculate the discrimination score of views and group views more reasonably. Results of the experiments show that the method can achieve SOTA performance. Sun et al. present a dual-matching method with cross-attention mechanism to address the limitations of matching-based methods caused by a preset fixed disparity range on depth estimation task. To tackle the mismatches on edges and details, they introduce an exquisite module based on left-right consistency. The method is proved to be competitive and effective by experiments conducted under popular benchmarks. Zhang et al. want to answer the following two questions: (1) does a query image with higher resolution than that of the gallery image also affect the pedestrian re-identification performance? If so, and (2) how does it affect performance? So, they propose an end-to-end trainable resolution independent person re-identification network that is composed of a cross-resolution Generative Adversarial Networks and embedding batch normalisation layers. The results demonstrate that the proposed method outperforms the SOTA methods in the pedestrian re-identification task on their expanded benchmark dataset. Huang et al. address the limitation of current gait-based age and gender recognition methods under multi-view scene, and propose an attention-aware spatio–temporal learning framework that employs silhouette sequence as an input to learn essential spatial–temporal gait representation. The proposed method has produced results that outperformed the benchmarks with an Mean Absolute Error of 6.68 years for age estimation and a Correct Classification Rate of 97% for gender classification. Zheng et al. apply deep learning to multi-view classroom behaviour detection. First, they propose an improved detection model based on YOLOv5 to improve the convergence speed of the prediction box. Second, they establish a quantitative evaluation standard for students' classroom attention, and then conduct training and verification by collecting multi-view classroom datasets. Finally, they increase the environment variation in the training model phase to make the model have better generalisation ability. Experiments demonstrate that the method can effectively identify and detect students' behaviours in the classroom from different views. Zhang et al. propose a method for multi-dimensional video anomaly detection, which uses the Object-meta instead of video frames as the input, and the Memory Search Guided Autoencoder with Memory Pools (MSGAE-MP) to reconstruct. The multi-dimensional information carried by the input can be strengthened via Object-meta. The MSGAE-MP construct multi-level memory pools, so as to reconstruct Object-meta in different dimensions. Experiments show that the method is feasible and has achieved excellent results. All of the papers published in this Special Issue show that multi-view representation learning theoretics have developed very fast in recent years. In addition, it is very promising to solve traditional computer vision tasks under multi-view setting, including but not limited to 3D object recognition, person re-identification, gait-based age and gender estimation, and depth estimation. Xin Ning and Chen Wang are responsible for the writing of Proposal and Editorial materials; Jun Zhou is responsible for the processing of articles; and Jing Wu, Lin Gu and Jian Cheng are responsible for the solicitation and publicity of the special issue. Firstly, we would like to thank all the authors for their innovative contributions and all the reviewers for their professional and crucial, yet constructive comments. Also, we wish to express our thanks to Mr Hang Ran, PhD students at Institute of Semiconductors, Chinese Academy of Sciences, for his assistance in this process. Last, we wish to express our gratitude to the editorial team of IET Computer Vision for their support throughout this venture. We hope you enjoy this collection of papers and that the Special Issue can stimulate further research and development in this area. This work is supported by the National Natural Science Foundation of China (Grant no. 61901436). National Natural Science Foundation of China, Grant/Award Number: 61901436. Data sharing is not applicable to this article as no new data were created or analyzed in this study. Xin Ning (SMIEEE) received a B.S. degree in software engineering in 2012, and a Ph.D. degree in electronic circuit and system from the university of Chinese Academy of Sciences, in 2017. He is currently an associate professor with the Laboratory of Artificial Neural Networks and High Speed Circuits, Institute of Semiconductors, Chinese Academy of Sciences. His current research interests include neural networks, intelligent systems and computer vision. He has published as the first or corresponding author in more than 45 papers in journals and refereed conferences. Now he serves as the young associated editor of CAAI Transactions on Intelligent Systems, the guest editor of Elsevier Journal on DISPLAYS. He is also the guest editor of CONNECTION SCIENCE and CONCURR COMP-PRACT E. He was the Website Chair of the IEEE HPBD&IS 2020 and the Publication Chair of the IEEE HPBD&IS 2021. Jun Zhou received a B.S. degree in computer science and a B.E. degree in international business from the Nanjing University of Science and Technology, Nanjing, China, in 1996 and 1998, respectively, an M.S. degree in computer science from Concordia University, Montreal, QC, Canada, in 2002, and a Ph.D. degree in computing science from the University of Alberta, Edmonton, AB, Canada, in 2006. He was a research fellow with the Research School of Computer Science, The Australian National University, Canberra, ACT, Australia, and a researcher with the Canberra Research Laboratory, National Information and Communications Technology Australia, Canberra. In 2012, he joined the School of Information and Communication Technology, Griffith University, Nathan, QLD, Australia, where he is currently a reader. His research interests include pattern recognition, computer vision, and spectral imaging and their applications in remote sensing and environmental informatics. He is the associate editor for the journal of Pattern Recognition and IEEE Trans. on Remote Sensing. Jian Cheng is a professor of Institute of Automation, Chinese Academy of Sciences. He received the B.S. and M.S. degrees in Mathematics from Wuhan University in 1998 and 2001, respectively. After that, he received a Ph.D degree in pattern recognition and intelligent systems from Institute of Automation, Chinese Academy of Sciences in 2004. His current major research interests include deep learning, computer vision, chip design, etc. Jing Wu is now a postdoc at the school of computer science, Beihang University. He received his B.E. degree from the school of computer science, Northwestern Polytechnical University in 2013 and received his PhD. degree from the school of computer science, Beihang University in 2021. His research interests include computer vision, stereo matching, 3D reconstruction and camera localization. Chen Wang is now a postdoc at the school of computer science, Beihang University. He received his B.E. degree from the school of computer science, Northwestern Polytechnical University in 2013 and received his PhD. degree from the school of computer science, Beihang University in 2021. His research interests include computer vision, stereo matching, 3D reconstruction and camera localization. Lin Gu received a B.Eng. degree from Shanghai University, Shanghai, China, in 2009, and a Ph.D. degree in computer vision from Australian National University in 2014. After Ph.D. graduation from the Australian National University, he worked as a post-doctoral researcher at A*STAR, Singapore. Then, he was a project researcher with the National Institute of Informatics, Japan, and also a visiting scholar with Kyoto University, Japan. He is currently a research scientist at RIKEN AIP, Japan, and a special researcher with the University of Tokyo, Japan. He is also an in-charge of a Moonshot and an ACT-X Project to improve artificial intelligence by simulating the human brain. His primary research interests lie in machine learning, medical imaging, and computational photography.
Xin Ning 0001, Jun Zhou 0001, Jian Cheng 0001, Jing Wu 0004, Chen Wang 0026, Lin Gu 0003
IET Comput. Vis.3
2025 One-Shot Cross-Domain Instance Detection With Universal Representation
Chen Feng 0002, Jian Cheng 0001, Yang Xiao 0007, Zhiguo Cao 0001
IEEE Trans Autom. Sci. Eng.2
2025 An Efficient Bit-Sparse DNN Accelerator Exploiting Adaptive Bit-Serial Computations
abstract
Bit sparsity, an intrinsic attribute of binary representation, has been widely utilized in DNN inference acceleration. Despite the advantages in performance and energy efficiency demonstrated by existing bit-serial-based bit-sparse accelerators, they still face two notable limitations: 1) At the low-level bit-serial multiplier level, existing methods either statically select weight or activation as the serialized object during the design phase, or simply serialize both without considering the distribution of non-zero bits in different operands, thereby failing to achieve optimal performance; 2) At the high-level dataflow level, existing approaches do not eliminate zero values in data movement and computation, leading to considerable energy and latency overhead, as well as suboptimal PE utilization. In this work, we propose AdaS-Pro accelerator for fast and energy-efficient DNN inference. At the multiplier level, AdaSPro employs an adaptive bit-serial computation scheme, which dynamically serializes the input operand with fewer non-zero bits at runtime, thereby minimizing compute cycles. To further enhance performance, AdaS-Pro introduces an improved Booth encoding method to reduce the number of non-zero bits in each operand. At the dataflow level, AdaS-Pro employs a compressed format to eliminate zero values and proposes a bi-directional inner-join unit coupled with a ring-shaped scheduler to achieve efficient non-zero workload extraction and balancing. Experimental results show that AdaS-Pro outperforms existing state-of-theart bit-sparse accelerators, such as BitLet, BitX, and Laconic, with performance improvements of 4.03×, 6.78×, and 1.43×, respectively.
Jiayao Ling, Gang Li 0015, Xiaolong Lin, Xing Li 0031, Jian Cheng 0001, Xiaoyao Liang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2024 Not All Tasks Are Equally Difficult: Multi-Task Deep Reinforcement Learning with Dynamic Depth Routing
abstract
Multi-task reinforcement learning endeavors to accomplish a set of different tasks with a single policy. To enhance data efficiency by sharing parameters across multiple tasks, a common practice segments the network into distinct modules and trains a routing network to recombine these modules into task-specific policies. However, existing routing approaches employ a fixed number of modules for all tasks, neglecting that tasks with varying difficulties commonly require varying amounts of knowledge. This work presents a Dynamic Depth Routing (D2R) framework, which learns strategic skipping of certain intermediate modules, thereby flexibly choosing different numbers of modules for each task. Under this framework, we further introduce a ResRouting method to address the issue of disparate routing paths between behavior and target policies during off-policy training. In addition, we design an automatic route-balancing mechanism to encourage continued routing exploration for unmastered tasks without disturbing the routing of mastered ones. We conduct extensive experiments on various robotics manipulation tasks in the Meta-World benchmark, where D2R achieves state-of-the-art performance with significantly improved learning efficiency.
Jinmin He, Kai Li 0022, Yifan Zang 0001, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
AAAI7
2024 Intrinsic Action Tendency Consistency for Cooperative Multi-Agent Reinforcement Learning
abstract
Efficient collaboration in the centralized training with decentralized execution (CTDE) paradigm remains a challenge in cooperative multi-agent systems. We identify divergent action tendencies among agents as a significant obstacle to CTDE's training efficiency, requiring a large number of training samples to achieve a unified consensus on agents' policies. This divergence stems from the lack of adequate team consensus-related guidance signals during credit assignment in CTDE. To address this, we propose Intrinsic Action Tendency Consistency, a novel approach for cooperative multi-agent reinforcement learning. It integrates intrinsic rewards, obtained through an action model, into a reward-additive CTDE (RA-CTDE) framework. We formulate an action model that enables surrounding agents to predict the central agent's action tendency. Leveraging these predictions, we compute a cooperative intrinsic reward that encourages agents to align their actions with their neighbors' predictions. We establish the equivalence between RA-CTDE and CTDE through theoretical analyses, demonstrating that CTDE's training process can be achieved using N individual targets. Building on this insight, we introduce a novel method to combine intrinsic rewards and RA-CTDE. Extensive experiments on challenging tasks in SMAC, MPE, and GRF benchmarks showcase the improved performance of our method.
Yifan Zhang 0001, Xi Sheryl Zhang, Yifan Zang 0001, Jian Cheng 0001
AAAI5
2024 Patch-Aware Sample Selection for Efficient Masked Image Modeling
abstract
Nowadays sample selection is drawing increasing attention. By extracting and training only on the most informative subset, sample selection can effectively reduce the training cost. Although sample selection is effective in conventional supervised learning, applying it to Masked Image Modeling (MIM) still poses challenges due to the gap between sample-level selection and patch-level pre-training. In this paper, we inspect the sample selection in MIM pre-training and find the basic selection suffers from performance degradation. We attribute this degradation primarily to 2 factors: the random mask strategy and the simple averaging function. We then propose Patch-Aware Sample Selection (PASS), including a low-cost Dynamic Trained Mask Predictor (DTMP) and Weighted Selection Score (WSS). DTMP consistently masks the informative patches in samples, ensuring a relatively accurate representation of selection score. WSS enhances the selection score using patch-level disparity. Extensive experiments show the effectiveness of PASS in selecting the most informative subset and accelerating pretraining. PASS exhibits superior performance across various datasets, MIM methods, and downstream tasks. Particularly, PASS improves MAE by 0.7% on ImageNet-1K while utilizing only 37% data budget and achieves ~1.7x speedup.
Zhengyang Zhuge, Yongjun Bao, Peisong Wang 0001, Jian Cheng 0001
AAAI6
2024 FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large Scale
abstract
Graph Neural Networks (GNNs) have shown great superiority on non-Euclidean graph data, achieving ground-breaking performance on various graph-related tasks. As a practical solution to train GNN on large graphs with billions of nodes and edges, the sampling-based training is widely adopted by existing training frameworks. However, through an in-depth analysis, we observe that the efficiency of existing sampling-based training frameworks is still limited due to the key bottlenecks lying in all three phases of sampling-based training, i.e., subgraph sample, memory IO, and computation. To this end, we propose FastGL, a GPU-efficient Framework for accelerating sampling-based training of GNN at Large scale by simultaneously optimizing all above three phases, taking into account both GPU characteristics and graph structure. Specifically, by exploiting the inherent overlap within graph structures, FastGL develops the Match-Reorder strategy to reduce the data traffic, which accelerates the memory IO without incurring any GPU memory overhead. Additionally, FastGL leverages a Memory-Aware computation method, harnessing the GPU memory's hierarchical nature to mitigate irregular data access during computation. FastGL further incorporates the Fused-Map approach aimed at diminishing the synchronization overhead during sampling. Extensive experiments demonstrate that FastGL can achieve an average speedup of 11.8×, 2.2× and 1.5× over the state-of-the-art frameworks PyG, DGL, and GNNLab, respectively. Our code is available at https://github.com/a1bc2def6g/fastgl-ae.
Peisong Wang 0001, Qinghao Hu 0001, Gang Li 0015, Xiaoyao Liang, Jian Cheng 0001
ASPLOS (4)6
2024 MEGA: A Memory-Efficient GNN Accelerator Exploiting Degree-Aware Mixed-Precision Quantization
abstract
Graph Neural Networks (GNNs) are becoming a promising technique in various domains due to their excellent capabilities in modeling non-Euclidean data. Although a spectrum of accelerators has been proposed to accelerate the inference of GNNs, our analysis demonstrates that the latency and energy consumption induced by DRAM access still significantly impedes the improvement of performance and energy efficiency. To address this issue, we propose a Memory - Efficient GNN Accelerator (MEGA) through algorithm and hardware co-design in this work. Specifically, at the algorithm level, through an in-depth analysis of the node property, we observe that the data-independent quantization in previous works is not optimal in terms of accuracy and memory efficiency. This motivates us to propose the Degree-Aware mixed-precision quantization method, in which a proper bitwidth is learned and allocated to a node according to its in-degree to compress GNNs as much as possible while maintaining accuracy. At the hardware level, we employ a heterogeneous architecture design in which the aggregation and combination phases are implemented separately with different dataflows. In order to boost the performance and energy efficiency, we also present an Adaptive-Package format to alleviate the storage overhead caused by the fine-grained bitwidth and diverse sparsity, and a Condense-Edge scheduling method to enhance the data locality and further alleviate the access irregularity induced by the extremely sparse adjacency matrix in the graph. We implement our MEGA accelerator in a 28nm technology node. Extensive experiments demonstrate that MEGA can achieve an average speedup of 38.3 ×, 7.1 ×, 4.0 ×, 3.6× and 47.6 ×, 7.2 ×, 5.4 ×, 4.5 × energy savings over four state-of-the-art GNN accelerators, HyGCN, GCNAX, GROW, and SGCN, respectively, while retaining task accuracy.
Fanrong Li, Gang Li 0015, Zejian Liu, Zitao Mo, Qinghao Hu 0001, Xiaoyao Liang, Jian Cheng 0001
HPCA8
2024 Towards Offline Opponent Modeling with In-context Learning
abstract
Opponent modeling aims at learning the opponent's behaviors, goals, or beliefs to reduce the uncertainty of the competitive environment and assist decision-making. Existing work has mostly focused on learning opponent models online, which is impractical and inefficient in practical scenarios. To this end, we formalize an Offline Opponent Modeling (OOM) problem with the objective of utilizing pre-collected offline datasets to learn opponent models that characterize the opponent from the viewpoint of the controlled agent, which aids in adapting to the unknown fixed policies of the opponent. Drawing on the promises of the Transformers for decision-making, we introduce a general approach, Transformer Against Opponent (TAO), for OOM. Essentially, TAO tackles the problem by harnessing the full potential of the supervised pre-trained Transformers' in-context learning capabilities. The foundation of TAO lies in three stages: an innovative offline policy embedding learning stage, an offline opponent-aware response policy training stage, and a deployment stage for opponent adaptation with in-context learning. Theoretical analysis establishes TAO's equivalence to Bayesian posterior sampling in opponent modeling and guarantees TAO's convergence in opponent policy recognition. Extensive experiments and ablation studies on competitive environments with sparse and dense rewards demonstrate the impressive performance of TAO. Our approach manifests remarkable prowess for fast adaptation, especially in the face of unseen opponent policies, confirming its in-context learning potency.
Yuheng Jing, Kai Li 0022, Bingyun Liu, Yifan Zang 0001, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
ICLR8
2024 Dynamic Discounted Counterfactual Regret Minimization
abstract
Counterfactual regret minimization (CFR) is a family of iterative algorithms showing promising results in solving imperfect-information games. Recent novel CFR variants (e.g., CFR+, DCFR) have significantly improved the convergence rate of the vanilla CFR. The key to these CFR variants’ performance is weighting each iteration non-uniformly, i.e., discounting earlier iterations. However, these algorithms use a fixed, manually-specified scheme to weight each iteration, which enormously limits their potential. In this work, we propose Dynamic Discounted CFR (DDCFR), the first equilibrium-finding framework that discounts prior iterations using a dynamic, automatically-learned scheme. We formalize CFR’s iteration process as a carefully designed Markov decision process and transform the discounting scheme learning problem into a policy optimization problem within it. The learned discounting scheme dynamically weights each iteration on the fly using information available at runtime. Experimental results across multiple games demonstrate that DDCFR’s dynamic discounting scheme has a strong generalization ability and leads to faster convergence with improved performance. The code is available at https://github.com/rpSebastian/DDCFR.
Hang Xu 0006, Kai Li 0022, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
ICLR6
2024 HGCN2SP: Hierarchical Graph Convolutional Network for Two-Stage Stochastic Programming
abstract
Two-stage Stochastic Programming (2SP) is a standard framework for modeling decision-making problems under uncertainty. While numerous methods exist, solving such problems with many scenarios remains challenging. Selecting representative scenarios is a practical method for accelerating solutions. However, current approaches typically rely on clustering or Monte Carlo sampling, failing to integrate scenario information deeply and overlooking the significant impact of the scenario order on solving time. To address these issues, we develop HGCN2SP, a novel model with a hierarchical graph designed for 2SP problems, encoding each scenario and modeling their relationships hierarchically. The model is trained in a reinforcement learning paradigm to utilize the feedback of the solver. The policy network is equipped with a hierarchical graph convolutional network for feature encoding and an attention-based decoder for scenario selection in proper order. Evaluation of two classic 2SP problems demonstrates that HGCN2SP provides high-quality decisions in a short computational time. Furthermore, HGCN2SP exhibits remarkable generalization capabilities in handling large-scale instances, even with a substantial number of variables or scenarios that were unseen during the training phase.
Yifan Zhang 0001, Zhenxing Liang, Jian Cheng 0001
ICML4
2024 Towards Efficient Spiking Transformer: a Token Sparsification Framework for Training and Inference Acceleration
abstract
Nowadays Spiking Transformers have exhibited remarkable performance close to Artificial Neural Networks (ANNs), while enjoying the inherent energy-efficiency of Spiking Neural Networks (SNNs). However, training Spiking Transformers on GPUs is considerably more time-consuming compared to the ANN counterparts, despite the energy-efficient inference through neuromorphic computation. In this paper, we investigate the token sparsification technique for efficient training of Spiking Transformer and find conventional methods suffer from noticeable performance degradation. We analyze the issue and propose our Sparsification with Timestep-wise Anchor Token and dual Alignments (STATA). Timestep-wise Anchor Token enables precise identification of important tokens across timesteps based on standardized criteria. Additionally, dual Alignments incorporate both Intra and Inter Alignment of the attention maps, fostering the learning of inferior attention. Extensive experiments show the effectiveness of STATA thoroughly, which demonstrates up to $\sim$1.53$\times$ training speedup and $\sim$48% energy reduction with comparable performance on various datasets and architectures.
Zhengyang Zhuge, Peisong Wang 0001, Xingting Yao, Jian Cheng 0001
ICML4
2024 Minimizing Weighted Counterfactual Regret with Optimistic Online Mirror Descent
Hang Xu 0006, Kai Li 0022, Bingyun Liu, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
IJCAI7
2024 Efficient Multi-task Reinforcement Learning with Cross-Task Policy Guidance
abstract
Multi-task reinforcement learning endeavors to efficiently leverage shared information across various tasks, facilitating the simultaneous learning of multiple tasks. Existing approaches primarily focus on parameter sharing with carefully designed network structures or tailored optimization procedures. However, they overlook a direct and complementary way to exploit cross-task similarities: the control policies of tasks already proficient in some skills can provide explicit guidance for unmastered tasks to accelerate skills acquisition. To this end, we present a novel framework called Cross-Task Policy Guidance (CTPG), which trains a guide policy for each task to select the behavior policy interacting with the environment from all tasks' control policies, generating better training trajectories. In addition, we propose two gating mechanisms to improve the learning efficiency of CTPG: one gate filters out control policies that are not beneficial for guidance, while the other gate blocks tasks that do not necessitate guidance. CTPG is a general framework adaptable to existing parameter sharing approaches. Empirical evaluations demonstrate that incorporating CTPG with these approaches significantly enhances performance in manipulation and locomotion benchmarks.
Jinmin He, Kai Li 0022, Yifan Zang 0001, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
NeurIPS7
2024 Opponent Modeling with In-context Search
abstract
Opponent modeling is a longstanding research topic aimed at enhancing decision-making by modeling information about opponents in multi-agent environments. However, existing approaches often face challenges such as having difficulty generalizing to unknown opponent policies and conducting unstable performance. To tackle these challenges, we propose a novel approach based on in-context learning and decision-time search named Opponent Modeling with In-context Search (OMIS). OMIS leverages in-context learning-based pretraining to train a Transformer model for decision-making. It consists of three in-context components: an actor learning best responses to opponent policies, an opponent imitator mimicking opponent actions, and a critic estimating state values. When testing in an environment that features unknown non-stationary opponent agents, OMIS uses pretrained in-context components for decision-time search to refine the actor's policy. Theoretically, we prove that under reasonable assumptions, OMIS without search converges in opponent policy recognition and has good generalization properties; with search, OMIS provides improvement guarantees, exhibiting performance stability. Empirically, in competitive, cooperative, and mixed environments, OMIS demonstrates more effective and stable adaptation to opponents than other approaches. See our project website at https://sites.google.com/view/nips2024-omis.
Yuheng Jing, Bingyun Liu, Kai Li 0022, Yifan Zang 0001, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
NeurIPS8
2024 GraphLeak: Patient Record Leakage through Gradients with Knowledge Graph
abstract
In real clinics, the medical data are scattered over multiple hospitals. Due to security and privacy concerns, it is almost impossible to gather all the data together and train a unified model. Therefore, multi-node machine learning systems are currently the mainstream form of model training in healthcare systems. Nevertheless, distributed training relies on the exchange of gradients, which has been proved under the risk of privacy leakage. That means malicious attackers can restore the user's sensitive data by utilizing the publicly shared gradients, which is a serious problem for extremely private data such as Electronic Healthcare Records (EHRs). The performance of the previous gradient attack method will drop rapidly when the batch size of training data increases, which makes it less threatening in practice. However, in this paper, we found in the medical domain, by leveraging prior knowledge like the medical knowledge graph, the leakage risk can be significantly amplified. In particular, we present GraphLeak, which incorporates the medical knowledge graph in gradient leakage attacks. GraphLeak can improve the restoration effect of gradient attacks even under large batches of data. We conduct experimental verification on electronic healthcare record datasets, including eICU and MIMIC-III. Our method has achieved state-of-the-art attack performance compared with previous works. Code is available at https://github.com/anonymous4ai/GraphLeak.
Xi Sheryl Zhang, Weifan Guan, Zhaopeng Qiu, Jian Cheng 0001, Xian Wu 0001, Yefeng Zheng 0001
WWW5
2024 Automatic labelling framework for optical remote sensing object detection samples in a wide area using deep learning
Ning Li 0033, Liang Cheng 0003, Hui Chen 0035, Yalu Zhang, Yunchang Yao, Jian Cheng 0001, Manchun Li 0004
Expert Syst. Appl.7
2024 Late better than early: A decision-level information fusion approach for RGB-Thermal crowd counting with illumination awareness
Jian Cheng 0001, Chen Feng 0002, Yang Xiao 0007, Zhiguo Cao 0001
Neurocomputing1
2024 Graph meets probabilistic generation model: A new perspective for graph disentanglement
Zouzhang Peng, Shuai Zheng 0005, Zhenfeng Zhu, Zhizhe Liu, Jian Cheng 0001, Honghui Dong, Yao Zhao 0001
Pattern Recognit.5
2024 Advance One-Shot Multispectral Instance Detection With Text's Supervision
abstract
One key issue within one-shot multispectral instance detection (OMID) is to extract features of strong instance discriminative power, domain adaptation capability, and instance-wise generality. Existing methods generally only rely on visual clues. Comparatively, text is advantageous due to its structured information, high semantics, and low noise. Inspired by recent emergence of large image-text datasets and breakthrough visual-language models, we propose to advance OMID with text's supervision for the first time. To this end, our key idea is to establish the relationship between one-shot multispectral instance with ImageNet class labels via the CLIP model. Particularly, we retrieve, rank, and ensemble the text features of ImageNet labels via instance image feature as query. Then the resulting instance image and text features are realigned and fused to obtain a multimodal feature. Meanwhile, a multispectral contrastive learning approach is proposed to drive multimodal feature learning for OMID. Note that all the procedures are end-to-end trained in a unified network. In this way, the instance discriminative power and domain adaptation capability are facilitated simultaneously. Experiments on two tailored multispectral instance detection datasets verify the effectiveness of our method.
Chen Feng 0002, Jian Cheng 0001, Yang Xiao 0007, Zhiguo Cao 0001
IEEE Signal Process. Lett.2
2024 Toward Accurate Binarized Neural Networks With Sparsity for Mobile Application
abstract
While binarized neural networks (BNNs) have attracted great interest, popular approaches proposed so far mainly exploit the symmetric sign function for feature binarization, i.e., to binarize activations into -1 and +1 with a fixed threshold of 0. However, whether this option is optimal has been largely overlooked. In this work, we propose the Sparsity-inducing BNN (Si-BNN) to quantize the activations to be either 0 or +1, which better approximates ReLU using 1-bit. We further introduce trainable thresholds into the backward function of binarization to guide the gradient propagation. Our method dramatically outperforms the current state-of-the-art, lowering the performance gap between full-precision networks and BNNs on mainstream architectures, achieving the new state-of-the-art on binarized AlexNet (Top-1 50.5%), ResNet-18 (Top-1 62.2%), and ResNet-50 (Top-1 68.3%). At inference time, Si-BNN still enjoys the high efficiency of bit-wise operations. In our implementation, the running time of binary AlexNet on the CPU can be competitive with the popular GPU-based deep learning framework.
Peisong Wang 0001, Jian Cheng 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 Asynchronous Event Processing with Local-Shift Graph Convolutional Network
abstract
Event cameras are bio-inspired sensors that produce sparse and asynchronous event streams instead of frame-based images at a high-rate. Recent works utilizing graph convolutional networks (GCNs) have achieved remarkable performance in recognition tasks, which model event stream as spatio-temporal graph. However, the computational mechanism of graph convolution introduces redundant computation when aggregating neighbor features, which limits the low-latency nature of the events. And they perform a synchronous inference process, which can not achieve a fast response to the asynchronous event signals. This paper proposes a local-shift graph convolutional network (LSNet), which utilizes a novel local-shift operation equipped with a local spatio-temporal attention component to achieve efficient and adaptive aggregation of neighbor features. To improve the efficiency of pooling operation in feature extraction, we design a node-importance based parallel pooling method (NIPooling) for sparse and low-latency event data. Based on the calculated importance of each node, NIPooling can efficiently obtain uniform sampling results in parallel, which retains the diversity of event streams. Furthermore, for achieving a fast response to asynchronous event signals, an asynchronous event processing procedure is proposed to restrict the network nodes which need to recompute activations only to those affected by the new arrival event. Experimental results show that the computational cost can be reduced by nearly 9 times through using local-shift operation and the proposed asynchronous procedure can further improve the inference efficiency, while achieving state-of-the-art performance on gesture recognition and object recognition.
Linhui Sun, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu
AAAI3
2023 TinyNeRF: Towards 100 x Compression of Voxel Radiance Fields
abstract
Voxel grid representation of 3D scene properties has been widely used to improve the training or rendering speed of the Neural Radiance Fields (NeRF) while at the same time achieving high synthesis quality. However, these methods accelerate the original NeRF at the expense of extra storage demand, which hinders their applications in many scenarios. To solve this limitation, we present TinyNeRF, a three-stage pipeline: frequency domain transformation, pruning and quantization that work together to reduce the storage demand of the voxel grids with little to no effects on their speed and synthesis quality. Based on the prior knowledge of visual signals sparsity in the frequency domain, we convert the original voxel grids in the frequency domain via block-wise discrete cosine transformation (DCT). Next, we apply pruning and quantization to enforce the DCT coefficients to be sparse and low-bit. Our method can be optimized from scratch in an end-to-end manner, and can typically compress the original models by 2 orders of magnitude with minimal sacrifice on speed and synthesis quality.
Tianli Zhao, Jiayuan Chen 0003, Cong Leng, Jian Cheng 0001
AAAI4
2023 TBERT: Dynamic BERT Inference with Top-k Based Predictors
abstract
Dynamic inference is a compression method that adaptively prunes unimportant components according to the input at the inference stage, which can achieve a better trade-off between computational complexity and model accuracy than static compression methods. However, there are two limitations in previous works. The first one is that they usually need to search the threshold on the evaluation dataset to achieve the target compression ratio, but the search process is non-trivial. The second one is that these methods are unstable. Their performance will be significantly degraded on some datasets, especially when the compression ratio is high. In this paper, we propose TBERT, a simple yet stable dynamic inference method. TBERT utilizes the top-k-based pruning strategy which allows accurate control of the compression ratio. To enable stable end-to-end training of the model, we carefully design the structure of the predictor. Moreover, we propose adding auxiliary classifiers to help the model's training. Experimental results on the GLUE benchmark demonstrate that our method achieves higher performance than previous state-of-the-art methods.
Zejian Liu, Jian Cheng 0001
DATE3
2023 On the Data-Efficiency with Contrastive Image Transformation in Reinforcement Learning
Xi Sheryl Zhang, Yushuo Li, Yifan Zhang 0001, Jian Cheng 0001
ICLR5
2023 $\rm A^2Q$: Aggregation-Aware Quantization for Graph Neural Networks
Fanrong Li, Zitao Mo, Qinghao Hu 0001, Gang Li 0015, Zejian Liu, Xiaoyao Liang, Jian Cheng 0001
ICLR8
2023 MCUNeRF: Packing NeRF into an MCU with 1MB Memory
abstract
Neural Radiance Fields (NeRFs) have revolutionized 3D scene synthesis. Voxel grids are commonly employed to enhance training or rendering speed, but they entail additional storage requirements. The large model size and high computational and memory demands impede their progress on resource-constrained devices, e.g., Microcontroller Units (MCUs). Besides, there is currently no NeRF rendering framework available on MCU devices. In this paper, we propose a NeRF method named MCUNeRF for 3D scene synthesis on MCU devices. The proposed MCUNeRF compresses voxel grids via a hybrid quantization algorithm merging learned step-size quantization (LSQ) and optimized product quantization (OPQ). To further reduce the model storage, we also propose a codebook-sharing method that renders multiple objects with a single quantization codebook. Then we implement a NeRF-based rendering framework for MCU devices, which leverages a low-bit neural network computation framework, i.e. CMSIS-NN, to accelerate the rendering progress. Extensive experiments on four datasets such as Synthetic-NeRF demonstrate that our proposed method could compress model data by 20-40 times with comparable rendering quality, which enables NeRF-based scene rendering on MCU devices with only 1M SRAM.
Zhixiang Ye, Qinghao Hu 0001, Tianli Zhao, Wangping Zhou, Jian Cheng 0001
ACM Multimedia5
2023 Towards Efficient and Accurate Winograd Convolution via Full Quantization
abstract
The Winograd algorithm is an efficient convolution implementation, which performs calculations in the transformed domain. To further improve the computation efficiency, recent works propose to combine it with model quantization. Although Post-Training Quantization has the advantage of low computational cost and has been successfully applied in many other scenarios, a severe accuracy drop exists when utilizing it in Winograd convolution. Besides, despite the Winograd algorithm consisting of four stages, most existing methods only quantize the element-wise multiplication stage, leaving a considerable portion of calculations in full precision. In this paper, observing the inconsistency among different transformation procedures, we present PTQ-Aware Winograd (PAW) to optimize them collaboratively under a unified objective function. Moreover, we explore the full quantization of faster Winograd (tile size $\geq4$) for the first time. We further propose a hardware-friendly method called Factorized Scale Quantization (FSQ), which can effectively balance the significant range differences in the Winograd domain. Experiments demonstrate the effectiveness of our method, e.g., with 8-bit quantization and a tile size of 6, our method outperforms the previous Winograd PTQ method by 8.27\% and 5.38\% in terms of the top-1 accuracy on ResNet-18 and ResNet-34, respectively.
Weihan Chen, Peisong Wang 0001, Jian Cheng 0001
NeurIPS5
2023 Automatic Grouping for Efficient Cooperative Multi-Agent Reinforcement Learning
abstract
Grouping is ubiquitous in natural systems and is essential for promoting efficiency in team coordination. This paper proposes a novel formulation of Group-oriented Multi-Agent Reinforcement Learning (GoMARL), which learns automatic grouping without domain knowledge for efficient cooperation. In contrast to existing approaches that attempt to directly learn the complex relationship between the joint action-values and individual utilities, we empower subgroups as a bridge to model the connection between small sets of agents and encourage cooperation among them, thereby improving the learning efficiency of the whole team. In particular, we factorize the joint action-values as a combination of group-wise values, which guide agents to improve their policies in a fine-grained fashion. We present an automatic grouping mechanism to generate dynamic groups and group action-values. We further introduce a hierarchical control for policy learning that drives the agents in the same group to specialize in similar policies and possess diverse strategies for various groups. Experiments on the StarCraft II micromanagement tasks and Google Research Football scenarios verify our method's effectiveness. Extensive component studies show how grouping works and enhances performance.
Yifan Zang 0001, Jinmin He, Kai Li 0022, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
NeurIPS7
2023 A regularization perspective based theoretical analysis for adversarial robustness of deep spiking neural networks
Hui Zhang 0098, Jian Cheng 0001, Jun Zhang 0024, Hongyi Liu 0001, Zhihui Wei
Neural Networks2
2023 Optimization-Based Post-Training Quantization With Bit-Split and Stitching
abstract
Deep neural networks have shown great promise in various domains. Meanwhile, problems including the storage and computing overheads arise along with these breakthroughs. To solve these problems, network quantization has received increasing attention due to its high efficiency and hardware-friendly property. Nonetheless, most existing quantization approaches rely on the full training dataset and the time-consuming fine-tuning process to retain accuracy. Post-training quantization does not have these problems, however, it has mainly been shown effective for 8-bit quantization. In this paper, we theoretically analyze the effect of network quantization and show that the quantization loss in the final output layer is bounded by the layer-wise activation reconstruction error. Based on this analysis, we propose an Optimization-based Post-training Quantization framework and a novel Bit-split optimization approach to achieve minimal accuracy degradation. The proposed framework is validated on a variety of computer vision tasks, including image classification, object detection, instance segmentation, with various network architectures. Specifically, we achieve near-original model performance even when quantizing FP32 models to 3-bit without fine-tuning.
Peisong Wang 0001, Weihan Chen, Qiang Chen 0007, Qingshan Liu 0001, Jian Cheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Towards Automatic Model Compression via a Unified Two-Stage Framework
Weihan Chen, Peisong Wang 0001, Jian Cheng 0001
Pattern Recognit.3
2023 Improving Extreme Low-Bit Quantization With Soft Threshold
abstract
Deep neural networks executing with low precision at inference time can gain acceleration and compression advantages over their high-precision counterparts, but need to overcome the challenge of accuracy degeneration as the bit-width decreases. This work focuses on under 4-bit quantization that has a significant accuracy degeneration. We start with ternarization, a balance between efficiency and accuracy that quantizes both weights and activations into ternary values. We find that the hard threshold$\Delta $introduced in previous ternary networks for determining quantization intervals and the suboptimal solution of$\Delta $limit the performance of the ternary model. To alleviate it, we present Soft Threshold Ternary Networks (STTN), which enables the model to automatically determine ternarized values instead of depending on a hard threshold. Based on it, we further generalize the idea of soft threshold from ternarization to arbitrary bit-width, named Soft Threshold Quantized Networks (STQN). We observe that previous quantization relies on the rounding-to-nearest function, constraining the quantization solution space and leading to a significant accuracy degradation, especially in low-bit ($\leq3$-bits) quantization. Instead of relying on the traditional rounding-to-nearest function, STQN is able to determine quantization intervals by itself adaptively. Accuracy experiments on image classification, object detection and instance segmentation, as well as efficiency experiments on field-programmable gate array (FPGA) demonstrate that the proposed framework can achieve a prominent tradeoff between accuracy and efficiency. Code is available at:https://github.com/WeixiangXu/STTN.
Fanrong Li, Yingying Jiang 0001, Yong A, Peisong Wang 0001, Jian Cheng 0001
IEEE Trans. Circuits Syst. Video Technol.7
2023 S&GDA: An Unsupervised Domain Adaptive Semantic Segmentation Framework Considering Both Imaging Scene and Geometric Domain Shifts
abstract
Unsupervised domain adaptation uses labeled data from a source domain to help learn a target domain without any labeled data. Previous studies have not systematically analyzed the causes of remote sensing (RS) domain shifts, making it difficult to effectively model domain shifts caused by differences in geographic scene and platform imaging positions and attitudes. Therefore, this study conducts detailed analysis of the causes of domain shifts in RS images, and an unsupervised domain adaptive semantic segmentation framework, called S&GDA, that considers both imaging scene and geometric domain shifts is proposed. S&GDA comprised two modules: imaging scene simulation and imaging geometric simulation modules. The imaging scene simulation module is instrumental in mitigating domain shifts in geographical scenes due to variations in natural and human factors, thereby achieving cross-domain imaging scene consistency. Meanwhile, the imaging geometric simulation module allows for accurate simulation of domain shifts caused by changes in the position and attitude of a platform, ensuring cross-domain imaging geometry consistency. Note that none of these modules add additional parameters or computational complexity to the model as they only work on the input side of the data. Comprehensive experiments are conducted on the LoveDA and ISPRS datasets to evaluate S&GDA. Results indicate that S&GDA outperforms the state-of-the-art (SOTA) unsupervised domain adaptive semantic segmentation method by 3.12% of mIoU and can achieve 85% of the performance of the fully supervised method.
Hui Chen 0035, Liang Cheng 0003, Ning Li 0033, Yunchang Yao, Jian Cheng 0001, Ka Zhang
IEEE Trans. Geosci. Remote. Sens.5
2023 ATF: An Alternating Training Framework for Weakly Supervised Face Alignment
abstract
In recent years, various face-landmark datasets have been published. Intuitively, it is significant to integrate multiple labeled datasets to achieve higher performance. Due to the different annotation schemes of datasets, it is hard to directly train models using them together. Although numerous efforts have been made in the joint use of datasets, there remain three shortages in previous methods,i.e., additional computation, limitation of the markups scheme, and limited support for the regression method. To solve the above issues, we proposed a novelAlternating Training Framework(ATF), which leverages the similarity and diversity across multiple datasets for a more robust detector. ATF mainly contains two sub-modules:Alternating Training with Decreasing Proportions(ATDP) andMixed Branch Loss($\mathcal {L}_{MB}$). In particular, ATDP trains multiple datasets simultaneously via a weakly supervised way to take advantage of the diversity among them, and$\mathcal {L}_{MB}$utilizes similar landmark pairs to constrain different branches of the corresponding datasets. Besides, we extend the framework to easily handle three situations: single target detector, joint detector, and novel detector. Extensive experiments demonstrate the effectiveness of our framework for both heatmap-based and direct coordinate regression. Moreover, we have achieved a joint detector that outperforms state-of-the-art methods on each benchmark.
Xing Lan, Qinghao Hu 0001, Jian Cheng 0001
IEEE Trans. Multim.3
2023 GPENs: Graph Data Learning With Graph Propagation-Embedding Networks
abstract
Compact representation of graph data is a fundamental problem in pattern recognition and machine learning area. Recently, graph neural networks (GNNs) have been widely studied for graph-structured data representation and learning tasks, such as graph semi-supervised learning, clustering, and low-dimensional embedding. In this article, we present graph propagation-embedding networks (GPENs), a new model for graph-structured data representation and learning problem. GPENs are mainly motivated by 1) revisiting of traditional graph propagation techniques for graph node context-aware feature representation and 2) recent studies on deeply graph embedding and neural network architecture. GPENs integrate both feature propagation on graph and low-dimensional embedding simultaneously into a unified network using a novel propagation-embedding architecture. GPENs have two main advantages. First, GPENs can be well-motivated and explained from feature propagation and deeply learning architecture. Second, the equilibrium representation of the propagation-embedding operation in GPENs has both exact and approximate formulations, both of which have simple closed-form solutions. This guarantees the compactivity and efficiency of GPENs. Third, GPENs can be naturally extended to multiple GPENs (M-GPENs) to address the data with multiple graph structures. Experiments on various semi-supervised learning tasks on several benchmark datasets demonstrate the effectiveness and benefits of the proposed GPENs and M-GPENs.
Bo Jiang 0002, Leiling Wang, Jian Cheng 0001, Jin Tang 0001, Bin Luo 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 Extremely Sparse Networks via Binary Augmented Pruning for Fast Image Classification
abstract
Network pruning and binarization have been demonstrated to be effective in neural network accelerator design for high speed and energy efficiency. However, most existing pruning approaches achieve a poor tradeoff between accuracy and efficiency, which on the other hand, has limited the progress of neural network accelerators. At the same time, binary networks are highly efficient, however, a large accuracy gap exists between binary networks and their full-precision counterparts. In this article, we investigate the merits of extremely sparse networks with binary connections for image classification through software-hardware codesign. More specifically, we first propose a binary augmented extremely pruning method that can achieve ~98% sparsity with small accuracy degradation. Then we design the hardware architecture based on the resulting sparse and binary networks, which extensively explores the benefits of extreme sparsity with negligible resource consumption introduced by binary branch. Experiments on large-scale ImageNet classification and field-programmable gate array (FPGA) demonstrate that the proposed software-hardware architecture can achieve a prominent tradeoff between accuracy and efficiency.
Peisong Wang 0001, Fanrong Li, Gang Li 0015, Jian Cheng 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 ECBC: Efficient Convolution via Blocked Columnizing
abstract
Direct convolution methods are now drawing increasing attention as they eliminate the additional storage demand required by indirect convolution algorithms (i.e., the transformed matrix generated by the im2col convolution algorithm). Nevertheless, the direct methods require special input-output tensor formatting, leading to extra time and memory consumption to get the desired data layout. In this article, we show that indirect convolution, if implemented properly, is able to achieve high computation performance with the help of highly optimized subroutines in matrix multiplication while avoid incurring substantial memory overhead. The proposed algorithm is called efficient convolution via blocked columnizing (ECBC). Inspired by the im2col convolution algorithm and the block algorithm of general matrix-to-matrix multiplication, we propose to conduct the convolution computation blockwisely. As a result, the tensor-to-matrix transformation process (e.g., the im2col operation) can also be done in a blockwise manner so that it only requires a small block of memory as small as the data block. Extensive experiments on various platforms and networks validate the effectiveness of ECBC, as well as the superiority of our proposed method against a set of widely used industrial-level convolution algorithms.
Tianli Zhao, Qinghao Hu 0001, Cong Leng, Jian Cheng 0001
IEEE Trans. Neural Networks Learn. Syst.7
2022 DPNAS: Neural Architecture Search for Deep Learning with Differential Privacy
abstract
Training deep neural networks (DNNs) for meaningful differential privacy (DP) guarantees severely degrades model utility. In this paper, we demonstrate that the architecture of DNNs has a significant impact on model utility in the context of private deep learning, whereas its effect is largely unexplored in previous studies. In light of this missing, we propose the very first framework that employs neural architecture search to automatic model design for private deep learning, dubbed as DPNAS. To integrate private learning with architecture search, a DP-aware approach is introduced for training candidate models composed on a delicately defined novel search space. We empirically certify the effectiveness of the proposed framework. The searched model DPNASNet achieves state-of-the-art privacy/utility trade-offs, e.g., for the privacy budget of (epsilon, delta)=(3, 1e-5), our model obtains test accuracy of 98.57% on MNIST, 88.09% on FashionMNIST, and 68.33% on CIFAR-10. Furthermore, by studying the generated architectures, we provide several intriguing findings of designing private-learning-friendly DNNs, which can shed new light on model design for deep learning with differential privacy.
Anda Cheng, Xi Sheryl Zhang, Qiang Chen 0007, Peisong Wang 0001, Jian Cheng 0001
AAAI6
2022 Towards Fully Sparse Training: Information Restoration with Spatial Similarity
abstract
The 2:4 structured sparsity pattern released by NVIDIA Ampere architecture, requiring four consecutive values containing at least two zeros, enables doubling math throughput for matrix multiplications. Recent works mainly focus on inference speedup via 2:4 sparsity while training acceleration has been largely overwhelmed where backpropagation consumes around 70% of the training time. However, unlike inference, training speedup with structured pruning is nontrivial due to the need to maintain the fidelity of gradients and reduce the additional overhead of performing 2:4 sparsity online. For the first time, this article proposes fully sparse training (FST) where `fully' indicates that ALL matrix multiplications in forward/backward propagation are structurally pruned while maintaining accuracy. To this end, we begin with saliency analysis, investigating the sensitivity of different sparse objects to structured pruning. Based on the observation of spatial similarity among activations, we propose pruning activations with fixed 2:4 masks. Moreover, an Information Restoration block is proposed to retrieve the lost information, which can be implemented by efficient gradient-shift operation. Evaluation of accuracy and efficiency shows that we can achieve 2× training acceleration with negligible accuracy degradation on challenging large-scale classification and detection tasks.
Ke Cheng 0002, Peisong Wang 0001, Jian Cheng 0001
AAAI5
2022 MixFormer: Mixing Features across Windows and Dimensions
abstract
While local-window self-attention performs notably in vision tasks, it suffers from limited receptive field and weak modeling capability issues. This is mainly because it performs self-attention within non-overlapped windows and shares weights on the channel dimension. We propose Mix-Former to find a solution. First, we combine local-window self-attention with depth-wise convolution in a parallel design, modeling cross-window connections to enlarge the receptive fields. Second, we propose bi-directional interactions across branches to provide complementary clues in the channel and spatial dimensions. These two designs are integrated to achieve efficient feature mixing among windows and dimensions. Our MixFormer provides competitive results on image classification with EfficientNet and shows better results than RegNet and Swin Transformer. Performance in downstream tasks outperforms its alternatives by significant margins with less computational costs in 5 dense prediction tasks on MS COCO, ADE20k, and LVIS. Code is available at https://github.com/PaddlePaddle/PaddleClas.
Qiang Chen 0007, Qiman Wu, Jian Wang 0066, Qinghao Hu 0001, Errui Ding, Jian Cheng 0001, Jingdong Wang 0001
CVPR7
2022 Differentially Private Federated Learning with Local Regularization and Sparsification
abstract
User-level differential privacy (DP) provides certifiable privacy guarantees to the information that is specific to any user's data in federated learning. Existing methods that ensure user-level DP come at the cost of severe accuracy decrease. In this paper, we study the cause of model performance degradation in federated learning with user-level DP guarantee. We find the key to solving this issue is to naturally restrict the norm of local updates before ex-ecuting operations that guarantee DP. To this end, we propose two techniques, Bounded Local Update Regularization and Local Update Sparsification, to increase model quality without sacrificing privacy. We provide theoretical analysis on the convergence of our framework and give rigorous privacy guarantees. Extensive experiments show that our framework significantly improves the privacy-utility trade-off over the state-of-the-arts for federated learning with user-level DP guarantee.
Anda Cheng, Peisong Wang 0001, Xi Sheryl Zhang, Jian Cheng 0001
CVPR4
2022 APRIL: Finding the Achilles' Heel on Privacy for Vision Transformers
abstract
Federated learning frameworks typically require collaborators to share their local gradient updates of a common model instead of sharing training data to preserve privacy. However, prior works on Gradient Leakage Attacks showed that private training data can be revealed from gradients. So far almost all relevant works base their attacks on fully-connected or convolutional neural networks. Given the recent overwhelmingly rising trend of adapting Transformers to solve multifarious vision tasks, it is highly valuable to investigate the privacy risk of vision transformers. In this paper, we analyse the gradient leakage risk of self-attention based mechanism in both theoretical and practical manners. Particularly, we propose APRIL - Attention PRIvacy Leakage, which poses a strong threat to self-attention inspired models such as ViT. Showing how vision Transformers are at the risk of privacy leakage via gradients, we urge the significance of designing privacy-safer Transformer models and defending schemes.
Xi Sheryl Zhang, Tianli Zhao, Jian Cheng 0001
CVPR5
2022 PalQuant: Accelerating High-Precision Networks on Low-Precision Accelerators
Qinghao Hu 0001, Gang Li 0015, Qiman Wu, Jian Cheng 0001
ECCV (11)4
2022 MENet: A Memory-Based Network with Dual-Branch for Efficient Event Stream Processing
Linhui Sun, Yifan Zhang 0001, Ke Cheng 0002, Jian Cheng 0001, Hanqing Lu
ECCV (24)4
2022 Multi-granularity Pruning for Model Acceleration on Mobile Devices
Tianli Zhao, Xi Sheryl Zhang, Jian Cheng 0001
ECCV (11)7
2022 Stacking More Linear Operations with Orthogonal Regularization to Learn Better
abstract
How to improve the generalization of CNN models has been a long-lasting problem in the deep learning community. This paper presents a runtime parameter/FLOPs-free method to strengthen CNN models by stacking linear convolution operations during training. We show that overparameterization with appropriate regularization can lead to a smooth optimization landscape that improves the performance. Concretely, we propose to add a 1 × 1 convolutional layer before and after the original k × k convolutional layer respectively, without any non-linear activations between them. In addition, Quasi-Orthogonal Regularization is proposed to maintain the added 1 × 1 filters as orthogonal matrixes. After training, those two 1 × 1 layers can be fused into the original k × k layer without changing the original network architecture, leaving no extra computations at inference, i.e. parameter/FLOPs-free.
Jian Cheng 0001
ICIP2
2022 Ristretto: An Atomized Processing Architecture for Sparsity-Condensed Stream Flow in CNN
abstract
Low-precision quantization and sparsity have been widely explored in CNN acceleration due to their effectiveness in reducing computational complexity and memory requirements. However, to support variable numerical precision and sparse computation, prior accelerators design flexible multipliers or sparse dataflow separately. A uniform solution that simultaneously exploits mixed-precision and dual-sided irregular sparsity for CNN acceleration is still lacking. Through an in-depth review of existing precision-scalable and sparse accelerators, we observe that a direct combination of low-level multipliers and high-level sparse dataflow from both sides is challenging due to their orthogonal design spaces. To this end, in this paper, we propose condensed streaming computation. By representing non-zero weights and activations as atomized streams, the low-level mixed-precision multiplication and high-level sparse convolution can be unified into a shared dataflow through hierarchical data reuse. Based on the condensed streaming computation, we propose Ristretto, an atomized architecture that exploits both mixed-precision and dual-sided irregular sparsity for CNN inference. We implement Ristretto in a 28nm technology node. Extensive evaluations show that Ristretto consistently outperforms three state-of-the-art CNN accelerators, including Bit Fusion, Laconic, and SparTen, in terms of performance and energy efficiency.
Gang Li 0015, Zhuoran Song, Naifeng Jing, Jian Cheng 0001, Xiaoyao Liang
MICRO5
2022 PKD: General Distillation Framework for Object Detectors via Pearson Correlation Coefficient
abstract
Knowledge distillation(KD) is a widely-used technique to train compact models in object detection. However, there is still a lack of study on how to distill between heterogeneous detectors. In this paper, we empirically find that better FPN features from a heterogeneous teacher detector can help the student although their detection heads and label assignments are different. However, directly aligning the feature maps to distill detectors suffers from two problems. First, the difference in feature magnitude between the teacher and the student could enforce overly strict constraints on the student. Second, the FPN stages and channels with large feature magnitude from the teacher model could dominate the gradient of distillation loss, which will overwhelm the effects of other features in KD and introduce much noise. To address the above issues, we propose to imitate features with Pearson Correlation Coefficient to focus on the relational information from the teacher and relax constraints on the magnitude of the features. Our method consistently outperforms the existing detection KD methods and works for both homogeneous and heterogeneous student-teacher pairs. Furthermore, it converges faster. With a powerful MaskRCNN-Swin detector as the teacher, ResNet-50 based RetinaNet and FCOS achieve 41.5% and 43.9% $mAP$ on COCO2017, which are 4.1% and 4.8% higher than the baseline, respectively.
Weihan Cao, Yifan Zhang 0001, Jianfei Gao 0003, Anda Cheng, Ke Cheng 0002, Jian Cheng 0001
NeurIPS6
2022 Singular Value Fine-tuning: Few-shot Segmentation requires Few-parameters Fine-tuning
abstract
Freezing the pre-trained backbone has become a standard paradigm to avoid overfitting in few-shot segmentation. In this paper, we rethink the paradigm and explore a new regime: {\em fine-tuning a small part of parameters in the backbone}. We present a solution to overcome the overfitting problem, leading to better model generalization on learning novel classes. Our method decomposes backbone parameters into three successive matrices via the Singular Value Decomposition (SVD), then {\em only fine-tunes the singular values} and keeps others frozen. The above design allows the model to adjust feature representations on novel classes while maintaining semantic clues within the pre-trained backbone. We evaluate our {\em Singular Value Fine-tuning (SVF)} approach on various few-shot segmentation methods with different backbones. We achieve state-of-the-art results on both Pascal-5$^i$ and COCO-20$^i$ across 1-shot and 5-shot settings. Hopefully, this simple baseline will encourage researchers to rethink the role of backbone fine-tuning in few-shot settings.
Yanpeng Sun, Qiang Chen 0007, Jian Wang 0066, Haocheng Feng, Junyu Han, Errui Ding, Jian Cheng 0001, Zechao Li, Jingdong Wang 0001
NeurIPS8
2022 GLIF: A Unified Gated Leaky Integrate-and-Fire Neuron for Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) have been studied over decades to incorporate their biological plausibility and leverage their promising energy efficiency. Throughout existing SNNs, the leaky integrate-and-fire (LIF) model is commonly adopted to formulate the spiking neuron and evolves into numerous variants with different biological features. However, most LIF-based neurons support only single biological feature in different neuronal behaviors, limiting their expressiveness and neuronal dynamic diversity. In this paper, we propose GLIF, a unified spiking neuron, to fuse different bio-features in different neuronal behaviors, enlarging the representation space of spiking neurons. In GLIF, gating factors, which are exploited to determine the proportion of the fused bio-features, are learnable during training. Combining all learnable membrane-related parameters, our method can make spiking neurons different and constantly changing, thus increasing the heterogeneity and adaptivity of spiking neurons. Extensive experiments on a variety of datasets demonstrate that our method obtains superior performance compared with other SNNs by simply changing their neuronal formulations to GLIF. In particular, we train a spiking ResNet-19 with GLIF and achieve $77.35\%$ top-1 accuracy with six time steps on CIFAR-100, which has advanced the state-of-the-art. Codes are available at https://github.com/Ikarosy/Gated-LIF.
Xingting Yao, Fanrong Li, Zitao Mo, Jian Cheng 0001
NeurIPS4
2022 FedFV: federated face verification via equivalent class embeddings
Lingyun Liu, Yifan Zhang 0001, Haoyuan Gao, Xingtao Yu, Jian Cheng 0001
Multim. Syst.5
2022 Action recognition via pose-based graph convolutional networks with intermediate dense supervision
Lei Shi 0018, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu
Pattern Recognit.3
2022 Block Convolution: Toward Memory-Efficient Inference of Large-Scale CNNs on FPGA
abstract
Deep convolutional neural networks have achieved remarkable progress in recent years. However, the large volume of intermediate results generated during inference poses a significant challenge to the accelerator design for resource-constrained field-programmable gate array (FPGA). Due to the limited on-chip storage, partial results of intermediate layers are frequently transferred back and forth between on-chip memory and off-chip DRAM, leading to a nonnegligible increase in latency and energy consumption. In this article, we propose block convolution, a hardware-friendly, simple, yet efficient convolution operation that can completely avoid the off-chip transfer of intermediate feature maps at runtime. The fundamental idea of block convolution is to eliminate the dependency of feature map tiles in the spatial dimension when spatial tiling is used, which is realized by splitting a feature map into independent blocks so that convolution can be performed separately on individual blocks. We conduct extensive experiments to demonstrate the efficacy of the proposed block convolution on both the algorithm side and the hardware side. Specifically, we evaluate block convolution on: 1) VGG-16, ResNet-18, ResNet-50, and MobileNet-V1 for the ImageNet classification task; 2) SSD and FPN for the COCO object detection task; and 3) VDSR for the Set5 single-image superresolution task. Experimental results demonstrate that comparable or higher accuracy can be achieved with block convolution. We also showcase two CNN accelerators via algorithm/hardware co-design based on block convolution on memory-limited FPGAs, and evaluation shows that both accelerators substantially outperform the baseline without off-chip transfer of intermediate feature maps.
Gang Li 0015, Zejian Liu, Fanrong Li, Jian Cheng 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 MFHI: Taking Modality-Free Human Identification as Zero-Shot Learning
abstract
Human identification is an important topic in event detection, person tracking, and public security. There have been numerous methods proposed for human identification, such as face identification, person re-identification, and gait identification. Typically, existing methods predominantly classify a queried image to a specific identity in an image gallery set (I2I). This is seriously limited for the scenario where only a textual description of the query or an attribute gallery set is available in a wide range of video surveillance applications (A2IorI2A). However, very few efforts have been devoted towards modality-free identification, i.e., identifying a query in a gallery set in a scalable way. In this work, we take an initial attempt, and formulate such a novelModality-FreeHumanIdentification (named MFHI) task as a generic zero-shot learning model in a scalable way. Meanwhile, it is capable of bridging the visual and semantic modalities by learning a discriminative prototype of each identity. In addition, the semantics-guided spatial attention is enforced on visual modality to obtain interpretable representations with both high global category-level and local attribute-level discrimination. Finally, we design and conduct an extensive group of experiments on two common challenging identification tasks, including face identification and person re-identification, demonstrating that our method outperforms a wide variety of state-of-the-art methods on modality-free human identification.
Zhizhe Liu, Xingxing Zhang 0001, Zhenfeng Zhu, Shuai Zheng 0005, Yao Zhao 0001, Jian Cheng 0001
IEEE Trans. Circuits Syst. Video Technol.6
2022 Adversarial Binary Mutual Learning for Semi-Supervised Deep Hashing
abstract
Hashing is a popular search algorithm for its compact binary representation and efficient Hamming distance calculation. Benefited from the advance of deep learning, deep hashing methods have achieved promising performance. However, those methods usually learn with expensive labeled data but fail to utilize unlabeled data. Furthermore, the traditional pairwise loss used by those methods cannot explicitly force similar/dissimilar pairs to small/large distances. Both weaknesses limit existing methods’ performance. To solve the first problem, we propose a novel semi-supervised deep hashing model named adversarial binary mutual learning (ABML). Specifically, our ABML consists of a generative model$G_{H}$and a discriminative model$D_{H}$, where$D_{H}$learns labeled data in a supervised way and$G_{H}$learns unlabeled data by synthesizing real images. We adopt an adversarial learning (AL) strategy to transfer the knowledge of unlabeled data to$D_{H}$by making$G_{H}$and$D_{H}$mutually learn from each other. To solve the second problem, we propose a novel Weibull cross-entropy loss (WCE) by using the Weibull distribution, which can distinguish tiny differences of distances and explicitly force similar/dissimilar distances as small/large as possible. Thus, the learned features are more discriminative. Finally, by incorporating ABML with WCE loss, our model can acquire more semantic and discriminative features. Extensive experiments on four common data sets (CIFAR-10, large database of handwritten digits (MNIST), ImageNet-10, and NUS-WIDE) and a large-scale data set ImageNet demonstrate that our approach successfully overcomes the two difficulties above and significantly outperforms state-of-the-art hashing methods.
Guan'an Wang, Qinghao Hu 0001, Yang Yang 0062, Jian Cheng 0001, Zeng-Guang Hou
IEEE Trans. Neural Networks Learn. Syst.4
2021 You Only Look One-Level Feature
abstract
This paper revisits feature pyramids networks (FPN) for one-stage detectors and points out that the success of FPN is due to its divide-and-conquer solution to the optimization problem in object detection rather than multi-scale feature fusion. From the perspective of optimization, we introduce an alternative way to address the problem instead of adopting the complex feature pyramids - utilizing only one-level feature for detection. Based on the simple and efficient solution, we present You Only Look One-level Feature (YOLOF). In our method, two key components, Dilated Encoder and Uniform Matching, are proposed and bring considerable improvements. Extensive experiments on the COCO benchmark prove the effectiveness of the proposed model. Our YOLOF achieves comparable results with its feature pyramids counterpart RetinaNet while being 2.5× faster. Without transformer layers, YOLOF can match the performance of DETR in a single-level feature manner with 7× less training epochs. Code is available at https://github.com/megvii-model/YOLOF.
Qiang Chen 0007, Tong Yang 0005, Xiangyu Zhang 0005, Jian Cheng 0001, Jian Sun 0001
CVPR5
2021 Hardware Acceleration of Fully Quantized BERT for Efficient Natural Language Processing
abstract
BERT is the most recent Transformer-based model that achieves state-of-the-art performance in various NLP tasks. In this paper, we investigate the hardware acceleration of BERT on FPGA for edge computing. To tackle the issue of huge computational complexity and memory footprint, we propose to fully quantize the BERT (FQ-BERT), including weights, activations, softmax, layer normalization, and all the intermediate results. Experiments demonstrate that the FQ-BERT can achieve 7.94× compression for weights with negligible performance loss. We then propose an accelerator tailored for the FQ-BERT and evaluate on Xilinx ZCU102 and ZCU11 FPGA. It can achieve a performance-per-watt of 3.18 fps/W, which is 28.91× and 12.72× over Intel(R) Core(TM) i7-8700 CPU and NVIDIA K80 GPU, respectively.
Zejian Liu, Gang Li 0015, Jian Cheng 0001
DATE3
2021 AdaSGN: Adapting Joint Number and Model Size for Efficient Skeleton-Based Action Recognition
abstract
Existing methods for skeleton-based action recognition mainly focus on improving the recognition accuracy, whereas the efficiency of the model is rarely considered. Recently, there are some works trying to speed up the skeleton modeling by designing light-weight modules. However, in addition to the model size, the amount of the data involved in the calculation is also an important factor for the running speed, especially for the skeleton data where most of the joints are redundant or non-informative to identify a specific skeleton. Besides, previous works usually employ one fix-sized model for all the samples regardless of the difficulty of recognition, which wastes computations for easy samples. To address these limitations, a novel approach, called AdaSGN, is proposed in this paper, which can reduce the computational cost of the inference process by adaptively controlling the input number of the joints of the skeleton on-the-fly. Moreover, it can also adaptively select the optimal model size for each sample to achieve a better trade-off between the accuracy and the efficiency. We conduct extensive experiments on three challenging datasets, namely, NTU-60, NTU-120 and SHREC, to verify the superiority of the proposed approach, where AdaSGN achieves comparable or even higher performance with much lower GFLOPs compared with the baseline method.
Lei Shi 0018, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu
ICCV3
2021 Towards Mixed-Precision Quantization of Neural Networks via Constrained Optimization
abstract
Quantization is a widely used technique to compress and accelerate deep neural networks. However, conventional quantization methods use the same bit-width for all (or most of) the layers, which often suffer significant accuracy degradation in the ultra-low precision regime and ignore the fact that emergent hardware accelerators begin to support mixed-precision computation. Consequently, we present a novel and principled framework to solve the mixed-precision quantization problem in this paper. Briefly speaking, we first formulate the mixed-precision quantization as a discrete constrained optimization problem. Then, to make the optimization tractable, we approximate the objective function with second-order Taylor expansion and propose an efficient approach to compute its Hessian matrix. Finally, based on the above simplification, we show that the original problem can be reformulated as a MultipleChoice Knapsack Problem (MCKP) and propose a greedy search algorithm to solve it efficiently. Compared with existing mixed-precision quantization works, our method is derived in a principled way and much more computationally efficient. Moreover, extensive experiments conducted on the ImageNet dataset and various kinds of network architectures also demonstrate its superiority over existing uniform and mixed-precision quantization approaches.
Weihan Chen, Peisong Wang 0001, Jian Cheng 0001
ICCV3
2021 Dynamic Dual Gating Neural Networks
abstract
In dynamic neural networks that adapt computations to different inputs, gating-based methods have demonstrated notable generality and applicability in trading-off the model complexity and accuracy. However, existing works only explore the redundancy from a single point of the network, limiting the performance. In this paper, we propose dual gating, a new dynamic computing method, to reduce the model complexity at run-time. For each convolutional block, dual gating identifies the informative features along two separate dimensions, spatial and channel. Specifically, the spatial gating module estimates which areas are essential, and the channel gating module predicts the salient channels that contribute more to the results. Then the computation of both unimportant regions and irrelevant channels can be skipped dynamically during inference. Extensive experiments on a variety of datasets demonstrate that our method can achieve higher accuracy under similar computing budgets compared with other dynamic execution methods. In particular, dynamic dual gating can provide 59.7% saving in computing of ResNet50 with 76.41% top-1 accuracy on ImageNet, which has advanced the state-of-the-art. Codes are available at https://github.com/lfr-0531/DGNet.
Fanrong Li, Gang Li 0015, Jian Cheng 0001
ICCV4
2021 Towards Binarized MobileNet via Structured Sparsity
Zhenmeng Zuo, Zhexin Li, Peisong Wang 0001, Weihan Chen, Jian Cheng 0001
ICIG (1)5
2021 3D-CenterNet: 3D object detection network for point clouds with center estimation priority
Qi Wang 0056, Jian Cheng 0001, Jianqiang Deng, Xinfang Zhang
Pattern Recognit.2
2021 SpatialFlow: Bridging All Tasks for Panoptic Segmentation
abstract
Object location is fundamental to panoptic segmentation as it is related to all things and stuff in the image scene. Knowing the locations of objects in the image provides clues for segmenting and helps the network better understand the scene. How to integrate object location in both thing and stuff segmentation is a crucial problem. In this article, we propose spatial information flows to achieve this objective. The flows can bridge all sub-tasks in panoptic segmentation by delivering the object's spatial context from the box regression task to others. More importantly, we design four parallel sub-networks to get a preferable adaptation of object spatial information in sub-tasks. Upon the sub-networks and the flows, we present a location-aware and unified framework for panoptic segmentation, denoted as SpatialFlow. We perform a detailed ablation study on each component and conduct extensive experiments to prove the effectiveness of SpatialFlow. Furthermore, we achieve state-of-the-art results, which are 47.9 PQ and 62.5 PQ respectively on MS-COCO and Cityscapes panoptic benchmarks.
Qiang Chen 0007, Anda Cheng, Peisong Wang 0001, Jian Cheng 0001
IEEE Trans. Circuits Syst. Video Technol.5
2021 Decoupled Two-Stage Crowd Counting and Beyond
abstract
One of appealing approaches to counting dense objects, such as crowd, is density map estimation. Density maps, however, present ambiguous appearance cues in congested scenes, rendering infeasibility in identifying individuals and difficulties in diagnosing errors. Inspired by an observation that counting can be interpreted as a two-stage process, i.e., identifying possible object regions and counting exact object numbers, we introduce a probabilistic intermediate representation termed the probability map that depicts the probability of each pixel being an object. This representation allows us to decouple counting into probability map regression (PMR) and count map regression (CMR). We therefore propose a novel decoupled two-stage counting (D2C) framework that sequentially regresses the probability map and learns a counter conditioned on the probability map. Given the probability map and the count map, a peak point detection algorithm is derived to localize each object with a point under the guidance of local counts. An advantage of D2C is that the counter can be learned reliably with additional synthesized probability maps. This addresses important data deficiency and sample imbalanced problems in counting. Our framework also enables easy diagnoses and analyses of error patterns. For instance, we find that, the counter per se is sufficiently accurate, while the bottleneck appears to be PMR. We further instantiate a network D2CNet in our framework and report state-of-the-art counting and localization performance across 6 crowd counting benchmarks. Since the probability map is a representation independent of visual appearance, D2CNet also exhibits remarkable cross-dataset transferability. Code and pretrained models are made available at: https://git.io/d2cnet.
Jian Cheng 0001, Haipeng Xiong, Zhiguo Cao 0001, Hao Lu 0003
IEEE Trans. Image Process.1
2021 Extremely Lightweight Skeleton-Based Action Recognition With ShiftGCN++
abstract
In skeleton-based action recognition, graph convolutional networks (GCNs) have achieved remarkable success. However, there are two shortcomings of current GCN-based methods. Firstly, the computation cost is pretty heavy, typically over 15 GFLOPs for one action sample. Some recent works even reach ~100 GFLOPs. Secondly, the receptive fields of both spatial graph and temporal graph are inflexible. Although recent works introduce incremental adaptive modules to enhance the expressiveness of spatial graph, their efficiency is still limited by regular GCN structures. In this paper, we propose a shift graph convolutional network (ShiftGCN) to overcome both shortcomings. ShiftGCN is composed of novel shift graph operations and lightweight point-wise convolutions, where the shift graph operations provide flexible receptive fields for both spatial graph and temporal graph. To further boost the efficiency, we introduce four techniques and build a more lightweight skeleton-based action recognition model named ShiftGCN++. ShiftGCN++ is an extremely computation-efficient model, which is designed for low-power and low-cost devices with very limited computing power. On three datasets for skeleton-based action recognition, ShiftGCN notably exceeds the state-of-the-art methods with over 10× less FLOPs and 4× practical speedup. ShiftGCN++ further boosts the efficiency of ShiftGCN, which achieves comparable performance with 6× less FLOPs and 2× practical speedup.
Ke Cheng 0002, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu
IEEE Trans. Image Process.4
2021 Unsupervised Network Quantization via Fixed-Point Factorization
abstract
The deep neural network (DNN) has achieved remarkable performance in a wide range of applications at the cost of huge memory and computational complexity. Fixed-point network quantization emerges as a popular acceleration and compression method but still suffers from huge performance degradation when extremely low-bit quantization is utilized. Moreover, current fixed-point quantization methods rely heavily on supervised retraining using large amounts of the labeled training data, while the labeled data are hard to obtain in the real-world applications. In this article, we propose an efficient framework, namely, fixed-point factorized network (FFN), to turn all weights into ternary values, i.e., {-1, 0, 1}. We highlight that the proposed FFN framework can achieve negligible degradation even without any supervised retraining on the labeled data. Note that the activations can be easily quantized into an 8-bit format; thus, the resulting networks only have low-bit fixed-point additions that are significantly more efficient than 32-bit floating-point multiply-accumulate operations (MACs). Extensive experiments on large-scale ImageNet classification and object detection on MS COCO show that the proposed FFN can achieve about more than 20× compression and remove most of the multiply operations with comparable accuracy. Codes are available on GitHub at https://github.com/wps712/FFN.
Peisong Wang 0001, Qiang Chen 0007, Anda Cheng, Qingshan Liu 0001, Jian Cheng 0001
IEEE Trans. Neural Networks Learn. Syst.6
2020 Sparsity-Inducing Binarized Neural Networks
abstract
Binarization of feature representation is critical for Binarized Neural Networks (BNNs). Currently, sign function is the commonly used method for feature binarization. Although it works well on small datasets, the performance on ImageNet remains unsatisfied. Previous methods mainly focus on minimizing quantization error, improving the training strategies and decomposing each convolution layer into several binary convolution modules. However, whether sign is the only option for binarization has been largely overlooked. In this work, we propose the Sparsity-inducing Binarized Neural Network (Si-BNN), to quantize the activations to be either 0 or +1, which introduces sparsity into binary representation. We further introduce trainable thresholds into the backward function of binarization to guide the gradient propagation. Our method dramatically outperforms current state-of-the-arts, lowering the performance gap between full-precision networks and BNNs on mainstream architectures, achieving the new state-of-the-art on binarized AlexNet (Top-1 50.5%), ResNet-18 (Top-1 59.7%), and VGG-Net (Top-1 63.2%). At inference time, Si-BNN still enjoys the high efficiency of exclusive-not-or (xnor) operations.
Peisong Wang 0001, Gang Li 0015, Tianli Zhao, Jian Cheng 0001
AAAI5
2020 M-NAS: Meta Neural Architecture Search
abstract
Neural Architecture Search (NAS) has recently outperformed hand-crafted networks in various areas. However, most prevalent NAS methods only focus on a pre-defined task. For a previously unseen task, the architecture is either searched from scratch, which is inefficient, or transferred from the one obtained on some other task, which might be sub-optimal. In this paper, we investigate a previously unexplored problem: whether a universal NAS method exists, such that task-aware architectures can be effectively generated? Towards this problem, we propose Meta Neural Architecture Search (M-NAS). To obtain task-specific architectures, M-NAS adopts a task-aware architecture controller for child model generation. Since optimal weights for different tasks and architectures span diversely, we resort to meta-learning, and learn meta-weights that efficiently adapt to a new task on the corresponding architecture with only several gradient descent steps. Experimental results demonstrate the superiority of M-NAS against a number of competitive baselines on both toy regression and few shot classification problems.
Jiaxiang Wu 0001, Haoli Bai, Jian Cheng 0001
AAAI4
2020 Cross-Modality Paired-Images Generation for RGB-Infrared Person Re-Identification
abstract
RGB-Infrared (IR) person re-identification is very challenging due to the large cross-modality variations between RGB and IR images. The key solution is to learn aligned features to the bridge RGB and IR modalities. However, due to the lack of correspondence labels between every pair of RGB and IR images, most methods try to alleviate the variations with set-level alignment by reducing the distance between the entire RGB and IR sets. However, this set-level alignment may lead to misalignment of some instances, which limits the performance for RGB-IR Re-ID. Different from existing methods, in this paper, we propose to generate cross-modality paired-images and perform both global set-level and fine-grained instance-level alignments. Our proposed method enjoys several merits. First, our method can perform set-level alignment by disentangling modality-specific and modality-invariant features. Compared with conventional methods, ours can explicitly remove the modality-specific features and the modality variation can be better reduced. Second, given cross-modality unpaired-images of a person, our method can generate cross-modality paired images from exchanged images. With them, we can directly perform instance-level alignment by minimizing distances of every pair of images. Extensive experimental results on two standard benchmarks demonstrate that the proposed model favourably against state-of-the-art methods. Especially, on SYSU-MM01 dataset, our model can achieve a gain of 9.2% and 7.7% in terms of Rank-1 and mAP. Code is available at https://github.com/wangguanan/JSIA-ReID.
Guan'an Wang, Tianzhu Zhang 0001, Yang Yang 0062, Jian Cheng 0001, Jianlong Chang, Zeng-Guang Hou
AAAI4
2020 Decoupled Spatial-Temporal Attention Network for Skeleton-Based Action-Gesture Recognition
Lei Shi 0018, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu
ACCV (5)3
2020 Towards Convolutional Neural Networks Compression via Global&Progressive Product Quantization
Weihan Chen, Peisong Wang 0001, Jian Cheng 0001
BMVC3
2020 Skeleton-Based Action Recognition With Shift Graph Convolutional Network
abstract
Action recognition with skeleton data is attracting more attention in computer vision. Recently, graph convolutional networks (GCNs), which model the human body skeletons as spatiotemporal graphs, have obtained remarkable performance. However, the computational complexity of GCN-based methods are pretty heavy, typically over 15 GFLOPs for one action sample. Recent works even reach about 100 GFLOPs. Another shortcoming is that the receptive fields of both spatial graph and temporal graph are inflexible. Although some works enhance the expressiveness of spatial graph by introducing incremental adaptive modules, their performance is still limited by regular GCN structures. In this paper, we propose a novel shift graph convolutional network (Shift-GCN) to overcome both shortcomings. Instead of using heavy regular graph convolutions, our Shift-GCN is composed of novel shift graph operations and lightweight point-wise convolutions, where the shift graph operations provide flexible receptive fields for both spatial graph and temporal graph. On three datasets for skeleton-based action recognition, the proposed Shift-GCN notably exceeds the state-of-the-art methods with more than 10 times less computational complexity.
Ke Cheng 0002, Yifan Zhang 0001, Weihan Chen, Jian Cheng 0001, Hanqing Lu
CVPR5
2020 Distribution-Induced Bidirectional Generative Adversarial Network for Graph Representation Learning
abstract
Graph representation learning aims to encode all nodes of a graph into low-dimensional vectors that will serve as input of many computer vision tasks. However, most existing algorithms ignore the existence of inherent data distribution and even noises. This may significantly increase the phenomenon of over-fitting and deteriorate the testing accuracy. In this paper, we propose a Distribution-induced Bidirectional Generative Adversarial Network (named DBGAN) for graph representation learning. Instead of the widely used Gaussian assumption, the prior distribution of latent representation in our DBGAN is estimated in a structure-aware way, which implicitly bridges the graph and content spaces by prototype learning. Thus discriminative and robust representations are generated for all nodes. Furthermore, to improve their generalization ability while preserving representation ability, the sample-level and distribution-level consistency are well balanced via a bidirectional adversarial learning framework. An extensive group of experiments is then carefully designed and presented, demonstrating that our DBGAN obtains remarkably more favorable trade-off between representation and robustness, and meanwhile is dimension-efficient, over currently available alternatives in various tasks.
Shuai Zheng 0005, Zhenfeng Zhu, Xingxing Zhang 0001, Zhizhe Liu, Jian Cheng 0001, Yao Zhao 0001
CVPR5
2020 Hardware Acceleration of CNN with One-Hot Quantization of Weights and Activations
abstract
In this paper, we propose a novel one-hot representation for weights and activations in CNN model and demonstrate its benefits on hardware accelerator design. Specifically, rather than merely reducing the bitwidth, we quantize both weights and activations into n-bit integers that containing only one non-zero bit per value. In this way, the massive multiply and accumulates (MACs) are equivalent to additions of powers of two that can be efficiently calculated with histogram based computaitons. Experiments on the ImageNet classification task show that comparable accuracy can be obtained on our proposed One-Hot Networks (OHN) compared to conventional fixed-point networks. As case studies, we evaluate the efficacy of the one-hot data representation on two state-of-the-art CNN accelerators on FPGA, our preliminary results show that 50% and 68.5% resource saving can be achieved on DaDianNao and Laconic respectively. Besides, the one-hot optimized Laconic can further achieve an average speedup of 4.94× on AlexNet.
Gang Li 0015, Peisong Wang 0001, Zejian Liu, Cong Leng, Jian Cheng 0001
DATE5
2020 Decoupling GCN with DropGraph Module for Skeleton-Based Action Recognition
Ke Cheng 0002, Yifan Zhang 0001, Congqi Cao, Lei Shi 0018, Jian Cheng 0001, Hanqing Lu
ECCV (24)5
2020 ProxyBNN: Learning Binarized Neural Networks via Proxy Matrices
Zitao Mo, Ke Cheng 0002, Qinghao Hu 0001, Peisong Wang 0001, Qingshan Liu 0001, Jian Cheng 0001
ECCV (3)8
2020 Faster Person Re-identification
Guan'an Wang, Shaogang Gong, Jian Cheng 0001, Zeng-Guang Hou
ECCV (8)3
2020 Rethinking The Pid Optimizer For Stochastic Optimization Of Deep Networks
abstract
Stochastic gradient descent with momentum (SGD-Momentum) always causes the overshoot problem due to the integral action of the momentum term. Recently, an ID optimizer is proposed to solve the overshoot problem with the help of derivative information. However, the derivative term suffers from the interference of the high-frequency noise, especially for the stochastic gradient descent method that uses minibatch data in each update step. In this work, we propose a complete PID optimizer, which weakens the effect of the D term and adds a P term to more stably alleviate the overshoot problem. To further reduce the interference of the high-frequency noise, two effective and efficient methods are proposed to stabilize the training process. Extensive experiments on three widely used benchmark datasets with different scales, i.e., MNIST, Cifar10 and TinyImageNet, demonstrate the superiority of our proposed PID optimizer on various popular deep neural networks.
Lei Shi 0018, Yifan Zhang 0001, Wanguo Wang, Jian Cheng 0001, Hanqing Lu
ICME4
2020 Towards Accurate Post-training Network Quantization via Bit-Split and Stitching
abstract
Network quantization is essential for deploying deep models to IoT devices due to its high efficiency. Most existing quantization approaches rely on the full training datasets and the time-consuming fine-tuning to retain accuracy. Post-training quantization does not have these problems, however, it has mainly been shown effective for 8-bit quantization due to the simple optimization strategy. In this paper, we propose a Bit-Split and Stitching framework (Bit-split) for lower-bit post-training quantization with minimal accuracy degradation. The proposed framework is validated on a variety of computer vision tasks, including image classification, object detection, instance segmentation, with various network architectures. Specifically, Bit-split can achieve near-original model performance even when quantizing FP32 models to INT3 without fine-tuning.
Peisong Wang 0001, Qiang Chen 0007, Jian Cheng 0001
ICML4
2020 Motion Complementary Network for Efficient Action Recognition
abstract
Both two-stream ConvNet and 3D ConvNet are widely used in action recognition. However, both methods are not efficient for deployment: calculating optical flow is very slow, while 3D convolution is computationally expensive. Our key insight is that the motion information from optical flow maps is complementary to the motion information from 3D ConvNet. Instead of simply combining these two methods, we propose two novel techniques to enhance the performance with less computational cost: fixed-motion-accumulation and balanced-motion-policy. With these two techniques, we propose a novel framework called Efficient Motion Complementary Network(EMC-Net) that enjoys both high efficiency and high performance. We conduct extensive experiments on Kinetics, UCF101, and Jester datasets. We achieve notably higher performance while consuming 4.7× less computation than I3D, 11.6× less computation than ECO, 17.8× less computation than R(2+1)D. On Kinetics dataset, we achieve 2.6% better performance than the recent proposed TSM with 1.4× fewer FLOPs and 10ms faster on K80 GPU.
Ke Cheng 0002, Yifan Zhang 0001, Chenghua Li, Jian Cheng 0001, Hanqing Lu
ICPR4
2020 PEAN: 3D Hand Pose Estimation Adversarial Network
abstract
Despite recent emerging research attention, 3D hand pose estimation still suffers from the problems of predicting inaccurate or invalid poses which conflict with physical and kinematic constraints. To address these problems, we propose a novel 3D hand pose estimation adversarial network (PEAN) which can implicitly utilize such constraints to regularize the prediction in an adversarial learning framework. PEAN contains two parts: a 3D hierarchical estimation network (3DHNet) to predict hand pose, which decouples the task into multiple subtasks with a hierarchical structure; a pose discrimination network (PDNet) to judge the reasonableness of the estimated 3D hand pose, which back-propagates the constraints to the estimation network. During the adversarial learning process, PDNet is expected to distinguish the estimated 3D hand pose and the ground truth, while 3DHNet is expected to estimate more valid pose to confuse PDNet. In this way, 3DHNet is capable of generating 3D poses with accurate positions and adaptively adjusting the invalid poses without additional prior knowledge. Experiments show that the proposed 3DHNet does a good job in predicting hand poses, and introducing PDNet to 3DHNet does further improve the accuracy and reasonableness of the predicted results. As a result, the proposed PEAN achieves the state-of-the-art performance on three public hand pose estimation datasets.
Linhui Sun, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu
ICPR3
2020 Soft Threshold Ternary Networks
abstract
Large neural networks are difficult to deploy on mobile devices because of intensive computation and storage. To alleviate it, we study ternarization, a balance between efficiency and accuracy that quantizes both weights and activations into ternary values. In previous ternarized neural networks, a hard threshold Δ is introduced to determine quantization intervals. Although the selection of Δ greatly affects the training results, previous works estimate Δ via an approximation or treat it as a hyper-parameter, which is suboptimal. In this paper, we present the Soft Threshold Ternary Networks (STTN), which enables the model to automatically determine quantization intervals instead of depending on a hard threshold. Concretely, we replace the original ternary kernel with the addition of two binary kernels at training time, where ternary values are determined by the combination of two corresponding binary values. At inference time, we add up the two binary kernels to obtain a single ternary kernel. Our method dramatically outperforms current state-of-the-arts, lowering the performance gap between full-precision networks and extreme low bit networks. Experiments on ImageNet with AlexNet (Top-1 55.6%), ResNet-18 (Top-1 66.2%) achieves new state-of-the-art.
Tianli Zhao, Qinghao Hu 0001, Peisong Wang 0001, Jian Cheng 0001
IJCAI6
2020 ATF: Towards Robust Face Alignment via Leveraging Similarity and Diversity across Different Datasets
abstract
Face alignment is an important task in the field of multi-media. Together with the impressive progress of algorithms, various benchmark datasets have been released in recent years. Intuitively, it is meaningful to integrate multiple labeled datasets with different annotations to achieve higher performance on a target landmark detector. Although numerous efforts have been made in joint usage, there yet remain three shortages in recent works, e.g., additional computation, limitation of the markups scheme, and limited support for the regression method. To address the above problems, we proposed a novel Alternating Training Framework (ATF), which leverages similarity and diversity across multi-media sources for a more robust detector. Our framework mainly contains two sub-modules: Alternating Training with Decreasing Proportions (ATDP) and Mixed Branch Loss (mathcal LMB). In particular, ATDP trains multiple datasets simultaneously to take advantage of the diversity between them, while mathcal LMB utilizes similar landmark pairs to constrain different branches of corresponding datasets. Extensive experiments on various benchmarks show the effectiveness of our framework, and ATF is feasible for both heatmap-based network and direct coordinate regression. Specifically, the mean error even reaches 3.17 on the experiment on 300W leveraging WFLW, which significantly outperforms state-of-the-art methods. Both in an ordinary convolutional network (OCN) and HRNET, ATF achieves up to 9.96% relative improvement. Our source codes are made publicly available at https://github.com/starhiking/ATF.
Xing Lan, Qinghao Hu 0001, Fangzhou Xiong, Cong Leng, Jian Cheng 0001
ACM Multimedia5
2020 Revisiting Parameter Sharing for Automatic Neural Channel Number Search
abstract
Recent advances in neural architecture search inspire many channel number search algorithms~(CNS) for convolutional neural networks. To improve searching efficiency, parameter sharing is widely applied, which reuses parameters among different channel configurations. Nevertheless, it is unclear how parameter sharing affects the searching process. In this paper, we aim at providing a better understanding and exploitation of parameter sharing for CNS. Specifically, we propose affine parameter sharing~(APS) as a general formulation to unify and quantitatively analyze existing channel search algorithms. It is found that with parameter sharing, weight updates of one architecture can simultaneously benefit other candidates. However, it also results in less confidence in choosing good architectures. We thus propose a new strategy of parameter sharing towards a better balance between training efficiency and architecture discrimination. Extensive analysis and experiments demonstrate the superiority of the proposed strategy in channel configuration against many state-of-the-art counterparts on benchmark datasets.
Haoli Bai, Jiaxiang Wu 0001, Xupeng Shi, Junzhou Huang, Irwin King, Michael R. Lyu, Jian Cheng 0001
NeurIPS8
2020 Convolutional prototype learning for zero-shot recognition
Zhizhe Liu, Xingxing Zhang 0001, Zhenfeng Zhu, Shuai Zheng 0005, Yao Zhao 0001, Jian Cheng 0001
Image Vis. Comput.6
2020 Cross-modality paired-images generation and augmentation for RGB-infrared person re-identification
Guan'an Wang, Yang Yang 0062, Tianzhu Zhang 0001, Jian Cheng 0001, Zeng-Guang Hou, Prayag Tiwari, Hari Mohan Pandey
Neural Networks4
2020 Robust one-stage object detection with location-aware classifiers
Qiang Chen 0007, Peisong Wang 0001, Anda Cheng, Wanguo Wang, Yifan Zhang 0001, Jian Cheng 0001
Pattern Recognit.6
2020 Gesture recognition based on deep deformable 3D convolutional neural networks
Yifan Zhang 0001, Lei Shi 0018, Yi Wu 0001, Ke Cheng 0002, Jian Cheng 0001, Hanqing Lu
Pattern Recognit.5
2020 FSA: A Fine-Grained Systolic Accelerator for Sparse CNNs
abstract
Sparsity, as an intrinsic property of convolutional neural networks (CNNs), has been widely employed for hardware acceleration, and many customized accelerators tailored for sparse weights or activations have been proposed in these years. However, the irregular sparse patterns introduced by both weights and activations are much more challenging for efficient computation. For example, due to the issues of access contention, workload imbalance, and tile fragmentation, the state-of-the-art sparse accelerator SCNN fails to fully leverage the benefits of sparsity, leading to nonoptimal results for both speedup and energy efficiency. In this article, we propose an efficient sparse CNN accelerator for both weights and activations, namely finegrained systolic accelerator (FSA), which jointly optimizes both hardware dataflow and software partitioning and scheduling strategy. Specifically, to deal with the access contentions problem, we present a fine-grained systolic dataflow, in which the activations move rhythmically along the horizontal processing element array while the weights are fed into the array in a fine-grained order. We then propose a hybrid network partitioning strategy that sets different partitioning strategies for different layers to balance the workload and alleviate the fragmentation problem caused by both sparse weights and activations. Finally, we present a scheduling search strategy to find the optimized schedules for neural networks, which can further improve energy efficiency. Extensive evaluations show that the proposed FSA consistently outperforms SCNN over AlexNet, VGGNet, GoogLeNet, and ResNet with an average speedup of 1.74× and up to 13.86× energy efficiency.
Fanrong Li, Gang Li 0015, Zitao Mo, Jian Cheng 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2020 Few-Shot Visual Classification Using Image Pairs With Binary Transformation
abstract
Accurately classifying images using few-shot samples have been widely explored by researchers. However, these methods have two drawbacks. First, images are often used independently. Second, class imbalance is ignored and hinders the classification accuracy with the increment of classes. To tackle these two drawbacks, in this paper, we propose a novel visual classification method using image pairs with binary transformation (IPBT). For one image, we bundle it with each training image into an image pair by concatenating the representations of the two images along with their similarity. The class consistency of two images is used to split the image pairs into binary groups. One group contains image pairs of the same class, while the other group consists of images pairs belonging to different classes. We train classifiers to separate the binary groups apart. To classify a testing image, we first bundle it with all the training images that are then predicted using the learned binary classifier. The image pair with the largest response is selected, and the testing image is assigned to the same class of the paired image. We conduct few-shot visual classification experiments on three public image datasets. The experimental results and analysis show the effectiveness of the proposed IPBT method.
Chunjie Zhang 0001, Chenghua Li, Jian Cheng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2020 A Quantum-Inspired Similarity Measure for the Analysis of Complete Weighted Graphs
abstract
We develop a novel method for measuring the similarity between complete weighted graphs, which are probed by means of the discrete-time quantum walks. Directly probing complete graphs using discrete-time quantum walks is intractable due to the cost of simulating the quantum walk. We overcome this problem by extracting a commute time minimum spanning tree from the complete weighted graph. The spanning tree is probed by a discrete-time quantum walk which is initialized using a weighted version of the Perron-Frobenius operator. This naturally encapsulates the edge weight information for the spanning tree extracted from the original graph. For each pair of complete weighted graphs to be compared, we simulate a discrete-time quantum walk on each of the corresponding commute time minimum spanning trees and, then, compute the associated density matrices for the quantum walks. The probability of the walk visiting each edge of the spanning tree is given by the diagonal elements of the density matrices. The similarity between each pair of graphs is then computed using either: 1) the inner product or 2) the negative exponential of the Jensen-Shannon divergence between the probability distributions. We show that in both cases the resulting similarity measure is positive definite and, therefore, corresponds to a kernel on the graphs. We perform a series of experiments on publicly available graph datasets from a variety of different domains, together with time-varying financial networks extracted from data for the New York Stock Exchange. Our experiments demonstrate the effectiveness of the proposed similarity measures.
Lu Bai 0001, Luca Rossi 0004, Lixin Cui, Jian Cheng 0001, Edwin R. Hancock
IEEE Trans. Cybern.4
2020 Multiview Semantic Representation for Visual Recognition
abstract
Due to interclass and intraclass variations, the images of different classes are often cluttered which makes it hard for efficient classifications. The use of discriminative classification algorithms helps to alleviate this problem. However, it is still an open problem to accurately model the relationships between visual representations and human perception. To alleviate these problems, in this paper, we propose a novel multiview semantic representation (MVSR) algorithm for efficient visual recognition. First, we leverage visually based methods to get initial image representations. We then use both visual and semantic similarities to divide images into groups which are then used for semantic representations. We treat different image representation strategies, partition methods, and numbers as different views. A graph is then used to combine the discriminative power of different views. The similarities between images can be obtained by measuring the similarities of graphs. Finally, we train classifiers to predict the categories of images. We evaluate the discriminative power of the proposed MVSR method for visual recognition on several public image datasets. Experimental results show the effectiveness of the proposed method.
Chunjie Zhang 0001, Jian Cheng 0001, Qi Tian 0001
IEEE Trans. Cybern.2
2020 Skeleton-Based Action Recognition With Multi-Stream Adaptive Graph Convolutional Networks
abstract
Graph convolutional networks (GCNs), which generalize CNNs to more generic non-Euclidean structures, have achieved remarkable performance for skeleton-based action recognition. However, there still exist several issues in the previous GCN-based models. First, the topology of the graph is set heuristically and fixed over all the model layers and input data. This may not be suitable for the hierarchy of the GCN model and the diversity of the data in action recognition tasks. Second, the second-order information of the skeleton data, i.e., the length and orientation of the bones, is rarely investigated, which is naturally more informative and discriminative for the human action recognition. In this work, we propose a novel multi-stream attention-enhanced adaptive graph convolutional neural network (MS-AAGCN) for skeleton-based action recognition. The graph topology in our model can be either uniformly or individually learned based on the input data in an end-to-end manner. This data-driven approach increases the flexibility of the model for graph construction and brings more generality to adapt to various data samples. Besides, the proposed adaptive graph convolutional layer is further enhanced by a spatial-temporal-channel attention module, which helps the model pay more attention to important joints, frames and features. Moreover, the information of both the joints and bones, together with their motion information, are simultaneously modeled in a multi-stream framework, which shows notable improvement for the recognition accuracy. Extensive experiments on the two large-scale datasets, NTU-RGBD and Kinetics-Skeleton, demonstrate that the performance of our model exceeds the state-of-the-art with a significant margin.
Lei Shi 0018, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu
IEEE Trans. Image Process.3
2020 Multi-View Image Classification With Visual, Semantic and View Consistency
abstract
Multi-view visual classification methods have been widely applied to use discriminative information of different views. This strategy has been proven very effective by many researchers. On the one hand, images are often treated independently without fully considering their visual and semantic correlations. On the other hand, view consistency is often ignored. To solve these problems, in this paper, we propose a novel multi-view image classification method with visual, semantic and view consistency (VSVC). For each image, we linearly combine multi-view information for image classification. The combination parameters are determined by considering both the classification loss and the visual, semantic and view consistency. Visual consistency is imposed by ensuring that visually similar images of the same view are predicted to have similar values. For semantic consistency, we impose the locality constraint that nearby images should be predicted to have the same class by multiview combination. View consistency is also used to ensure that similar images have consistent multi-view combination parameters. An alternative optimization strategy is used to learn the combination parameters. To evaluate the effectiveness of VSVC, we perform image classification experiments on several public datasets. The experimental results on these datasets show the effectiveness of the proposed VSVC method.
Chunjie Zhang 0001, Jian Cheng 0001, Qi Tian 0001
IEEE Trans. Image Process.2
2019 ODE-Inspired Network Design for Single Image Super-Resolution
abstract
Single image super-resolution, as a high dimensional structured prediction problem, aims to characterize fine-grain information given a low-resolution sample. Recent advances in convolutional neural networks are introduced into super-resolution and push forward progress in this field. Current studies have achieved impressive performance by manually designing deep residual neural networks but overly relies on practical experience. In this paper, we propose to adopt an ordinary differential equation (ODE)-inspired design scheme for single image super-resolution, which have brought us a new understanding of ResNet in classification problems. Not only is it interpretable for super-resolution but it provides a reliable guideline on network designs. By casting the numerical schemes in ODE as blueprints, we derive two types of network structures: LF-block and RK-block, which correspond to the Leapfrog method and Runge-Kutta method in numerical ordinary differential equations. We evaluate our models on benchmark datasets, and the results show that our methods surpass the state-of-the-arts while keeping comparable parameters and operations.
Zitao Mo, Peisong Wang 0001, Yang Liu 0021, Mingyuan Yang, Jian Cheng 0001
CVPR6
2019 K-Nearest Neighbors Hashing
abstract
Hashing based approximate nearest neighbor search embeds high dimensional data to compact binary codes, which enables efficient similarity search and storage. However, the non-isometry sign(·) function makes it hard to project the nearest neighbors in continuous data space into the closest codewords in discrete Hamming space. In this work, we revisit the sign(·) function from the perspective of space partitioning. In specific, we bridge the gap between k-nearest neighbors and binary hashing codes with Shannon entropy. We further propose a novel K-Nearest Neighbors Hashing (KNNH) method to learn binary representations from KNN within the subspaces generated by sign(·). Theoretical and experimental results show that the KNN relation is of central importance to neighbor preserving embeddings, and the proposed method outperforms the state-of-the-arts on benchmark datasets.
Peisong Wang 0001, Jian Cheng 0001
CVPR3
2019 Skeleton-Based Action Recognition With Directed Graph Neural Networks
abstract
The skeleton data have been widely used for the action recognition tasks since they can robustly accommodate dynamic circumstances and complex backgrounds. In existing methods, both the joint and bone information in skeleton data have been proved to be of great help for action recognition tasks. However, how to incorporate these two types of data to best take advantage of the relationship between joints and bones remains a problem to be solved. In this work, we represent the skeleton data as a directed acyclic graph based on the kinematic dependency between the joints and bones in the natural human body. A novel directed graph neural network is designed specially to extract the information of joints, bones and their relations and make prediction based on the extracted features. In addition, to better fit the action recognition task, the topological structure of the graph is made adaptive based on the training process, which brings notable improvement. Moreover, the motion information of the skeleton sequence is exploited and combined with the spatial information to further enhance the performance in a two-stream framework. Our final model is tested on two large-scale datasets, NTU-RGBD and Skeleton-Kinetics, and exceeds state-of-the-art performance on both of them.
Lei Shi 0018, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu
CVPR3
2019 Two-Stream Adaptive Graph Convolutional Networks for Skeleton-Based Action Recognition
abstract
In skeleton-based action recognition, graph convolutional networks (GCNs), which model the human body skeletons as spatiotemporal graphs, have achieved remarkable performance. However, in existing GCN-based methods, the topology of the graph is set manually, and it is fixed over all layers and input samples. This may not be optimal for the hierarchical GCN and diverse samples in action recognition tasks. In addition, the second-order information (the lengths and directions of bones) of the skeleton data, which is naturally more informative and discriminative for action recognition, is rarely investigated in existing methods. In this work, we propose a novel two-stream adaptive graph convolutional network (2s-AGCN) for skeleton-based action recognition. The topology of the graph in our model can be either uniformly or individually learned by the BP algorithm in an end-to-end manner. This data-driven method increases the flexibility of the model for graph construction and brings more generality to adapt to various data samples. Moreover, a two-stream framework is proposed to model both the first-order and the second-order information simultaneously, which shows notable improvement for the recognition accuracy. Extensive experiments on the two large-scale datasets, NTU-RGBD and Kinetics-Skeleton, demonstrate that the performance of our model exceeds the state-of-the-art with a significant margin.
Lei Shi 0018, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu
CVPR3
2019 RGB-Infrared Cross-Modality Person Re-Identification via Joint Pixel and Feature Alignment
abstract
RGB-Infrared (IR) person re-identification is an important and challenging task due to large cross-modality variations between RGB and IR images. Most conventional approaches aim to bridge the cross-modality gap with feature alignment by feature representation learning. Different from existing methods, in this paper, we propose a novel and end-to-end Alignment Generative Adversarial Network (AlignGAN) for the RGB-IR RE-ID task. The proposed model enjoys several merits. First, it can exploit pixel alignment and feature alignment jointly. To the best of our knowledge, this is the first work to model the two alignment strategies jointly for the RGB-IR RE-ID problem. Second, the proposed model consists of a pixel generator, a feature generator and a joint discriminator. By playing a min-max game among the three components, our model is able to not only alleviate the cross-modality and intra-modality variations, but also learn identity-consistent features. Extensive experimental results on two standard benchmarks demonstrate that the proposed model performs favourably against state-of-the-art methods. Especially, on SYSU-MM01 dataset, our model can achieve an absolute gain of 15.4% and 12.9% in terms of Rank-1 and mAP.
Guan'an Wang, Tianzhu Zhang 0001, Jian Cheng 0001, Si Liu 0001, Yang Yang 0062, Zeng-Guang Hou
ICCV3
2019 DeLTR: A Deep Learning Based Approach to Traffic Light Recognition
Yiyang Cai, Chenghua Li, Sujuan Wang, Jian Cheng 0001
ICIG (3)4
2019 Gesture Recognition Using Spatiotemporal Deformable Convolutional Representation
abstract
Dynamic gesture recognition, which plays an essential role in human-computer interaction, has been widely investigated but not yet addressed. The interference of the varied and complex background makes the classifier easily be misguided due to the relatively smaller size of the hands and arms compared with the full scenes. In this paper, we address the problem by proposing a novel spatiotemporal deformable convolutional neural network for end-to-end learning. To eliminate the background interference, a light-weight spatiotemporal deformable convolution module is specially designed to augment the spatiotemporal sampling locations of 3D convolution by learning additional offsets according to the preceding feature map. The proposed method is evaluated on two challenging datasets, EgoGesture and Jester, and achieves the state-of-the-art performance on both of the two datasets. The code and trained models will be released for better communication and future work.
Lei Shi 0018, Yifan Zhang 0001, Jian Cheng 0001, Hanqing Lu
ICIP4
2019 Reading selectively via Binary Input Gated Recurrent Unit
abstract
Recurrent Neural Networks (RNNs) have shown great promise in sequence modeling tasks. Gated Recurrent Unit (GRU) is one of the most used recurrent structures, which makes a good trade-off between performance and time spent. However, its practical implementation based on soft gates only partially achieves the goal to control information flow. We can hardly explain what the network has learnt internally. Inspired by human reading, we introduce binary input gated recurrent unit (BIGRU), a GRU based model using a binary input gate instead of the reset gate in GRU. By doing so, our model can read selectively during interference. In our experiments, we show that BIGRU mainly ignores the conjunctions, adverbs and articles that do not make a big difference to the document understanding, which is meaningful for us to further understand how the network works. In addition, due to reduced interference from redundant information, our model achieves better performances than baseline GRU in all the testing tasks.
Peisong Wang 0001, Hanqing Lu, Jian Cheng 0001
IJCAI4
2019 Color-Sensitive Person Re-Identification
abstract
Recent deep Re-ID models mainly focus on learning high-level semantic features, while failing to explicitly explore color information which is one of the most important cues for person Re-ID. In this paper, we propose a novel Color-Sensitive Re-ID to take full advantage of color information. On one hand, we train our model with real and fake images. By using the extra fake images, more color information can be exploited and it can avoid overfitting during training. On the other hand, we also train our model with images of the same person with different colors. By doing so, features can be forced to focus on the color difference in regions. To generate fake images with specified colors, we propose a novel Color Translation GAN (CTGAN) to learn mappings between different clothing colors and preserve identity consistency among the same clothing color. Extensive evaluations on two benchmark datasets show that our approach significantly outperforms state-of-the-art Re-ID models.
Guan'an Wang, Yang Yang 0062, Jian Cheng 0001, Jinqiao Wang, Zeng-Guang Hou
IJCAI3
2019 Diffusion induced graph representation learning
Fuzhen Li, Zhenfeng Zhu, Xingxing Zhang 0001, Jian Cheng 0001, Yao Zhao 0001
Neurocomputing4
2019 BitStream: An efficient framework for inference of binary neural networks on CPUs
Yanshu Jiang, Tianli Zhao, Cong Leng, Jian Cheng 0001
Pattern Recognit. Lett.5
2019 Edge Heuristic GAN for Non-Uniform Blind Deblurring
abstract
Non-uniform blur, mainly caused by camera shake and motions of multiple objects, is one of the most common causes of image quality degradation. However, the traditional blind deblurring methods based on blur kernel estimation do not perform well on complicated non-uniform motion blurs. However, recent studies show that GAN-based approaches achieve impressive performance on deblurring tasks. In this letter, to further improve the performance of GAN-based methods on deblurring tasks, we propose an edge heuristic multi-scale generative adversarial network (GAN), which uses the coarse-to-fine scheme to restore clear images in an end-to-end manner. In particular, an edge-generated network is designed to generate sharp edges as auxiliary information to guide the deblurring process. Furthermore, We propose a hierarchical content loss function for deblurring tasks. Extensive experiments on different datasets show that our method achieves state-of-the-art performance in dynamic scene deblurring.
Shuai Zheng 0005, Zhenfeng Zhu, Jian Cheng 0001, Yandong Guo, Yao Zhao 0001
IEEE Signal Process. Lett.3
2019 Multiview, Few-Labeled Object Categorization by Predicting Labels With View Consistency
abstract
The categorization accuracies of objects have been greatly improved in recent years. However, large quantities of labeled images are needed. many methods fail when only few labeled images are available. To tackle the few-labeled object categorization problem, we need to represent and classify them from multiple views. In this paper, we propose a novel multiview, few-labeled object categorization algorithm by predicting the labels of images with view consistency (MVFL-VC). We use labeled images along with other unlabeled images in a unified framework. A mapping function is learned to model the correlations of images with their labels. Since there are no labeling information for unlabeled images, we simultaneously learn the mapping function and image labels by classification error minimization. We make use of multiview information for joint object categorization. Although different views represent different aspects of images, for one image, the predicted categories of multiple views should be consistent with each other. We learn the mapping function by minimizing the summed classification losses along with the discrepancy of predicted labels between different views in an alternative way. We conduct object categorization experiments on five public image datasets and compare with other semi-supervised methods. Experimental results well demonstrate the effectiveness of the proposed MVFL-VC method.
Chunjie Zhang 0001, Jian Cheng 0001, Qi Tian 0001
IEEE Trans. Cybern.2
2019 Unsupervised and Semi-Supervised Image Classification With Weak Semantic Consistency
abstract
Supervised methods have been widely used for image classifications. Although great progress has been made, existing supervised methods rely on well-labeled samples for classification. However, we often have large quantities of images with few or no labels. To cope with this problem, in this paper, we propose a novel weak semantic consistency constrained image classification method. We start from an extreme circumstance by viewing each image as one class. We train exemplar classifiers to separate each image from other images. For each image, we use the learned exemplar classifiers to predict the weak semantic correlations with the exemplar classifiers. When no labeled information is available, we cluster images using the weak semantic correlations and assign images within one cluster to the same mid-level class. When partially labeled images are available, we can use them to constrain the clustering process by assigning images of varied semantics to different mid-level classes. We use the newly assigned images for classifier training and new image representations, which can then be used for similar image assignments. The classifier training, image representation, and assignment processes are repeated until convergence. We conduct both unsupervised and semi-supervised image classification experiments on several datasets. The experimental results show the effectiveness of the proposed unsupervised and semi-supervised weak semantic consistency image classification method.
Chunjie Zhang 0001, Jian Cheng 0001, Qi Tian 0001
IEEE Trans. Multim.2
2019 Semantically Modeling of Object and Context for Categorization
abstract
Object-centric-based categorization methods have been proven more effective than hard partitions of images (e.g., spatial pyramid matching). However, how to determine the locations of objects is still an open problem. Besides, modeling of context areas is often mixed with the background. Moreover, the semantic information is often ignored by these methods that only use visual representations for classification. In this paper, we propose an object categorization method by semantically modeling the object and context information (SOC). We first select a number of candidate regions with high confidence scores and semantically represent these regions by measuring correlations of each region with prelearned classifiers (e.g., local feature-based classifiers and deep convolutional-neural-network-based classifiers). These regions are clustered for object selections. The other selected areas are then viewed as context areas. We treat other areas beyond the object and context areas within one image as the background. The visually and semantically represented objects and contexts are then used along with the background area for object representations and categorizations. Experimental results on several public data sets well demonstrate the effectiveness of the proposed object categorization method by semantically modeling the object and context information.
Chunjie Zhang 0001, Jian Cheng 0001, Qi Tian 0001
IEEE Trans. Neural Networks Learn. Syst.2
2018 From Hashing to CNNs: Training Binary Weight Networks via Hashing
abstract
Deep convolutional neural networks (CNNs) have shown appealing performance on various computer vision tasks in recent years. This motivates people to deploy CNNs to real-world applications. However, most of state-of-art CNNs require large memory and computational resources, which hinders the deployment on mobile devices. Recent studies show that low-bit weight representation can reduce much storage and memory demand, and also can achieve efficient network inference. To achieve this goal, we propose a novel approach named BWNH to train Binary Weight Networks via Hashing. In this paper, we first reveal the strong connection between inner-product preserving hashing and binary weight networks, and show that training binary weight networks can be intrinsically regarded as a hashing problem. Based on this perspective, we propose an alternating optimization method to learn the hash codes instead of directly learning binary weights. Extensive experiments on CIFAR10, CIFAR100 and ImageNet demonstrate that our proposed BWNH outperforms current state-of-art by a large margin.
Qinghao Hu 0001, Peisong Wang 0001, Jian Cheng 0001
AAAI3
2018 Two-Step Quantization for Low-Bit Neural Networks
abstract
Every bit matters in the hardware design of quantized neural networks. However, extremely-low-bit representation usually causes large accuracy drop. Thus, how to train extremely-low-bit neural networks with high accuracy is of central importance. Most existing network quantization approaches learn transformations (low-bit weights) as well as encodings (low-bit activations) simultaneously. This tight coupling makes the optimization problem difficult, and thus prevents the network from learning optimal representations. In this paper, we propose a simple yet effective Two-Step Quantization (TSQ) framework, by decomposing the network quantization problem into two steps: code learning and transformation function learning based on the learned codes. For the first step, we propose the sparse quantization method for code learning. The second step can be formulated as a non-linear least square regression problem with low-bit constraints, which can be solved efficiently in an iterative manner. Extensive experiments on CIFAR-10 and ILSVRC-12 datasets demonstrate that the proposed TSQ is effective and outperforms the state-of-the-art by a large margin. Especially, for 2-bit activation and ternary weight quantization of AlexNet, the accuracy of our TSQ drops only about 0.5 points compared with the full-precision counterpart, outperforming current state-of-the-art by more than 5 points.
Peisong Wang 0001, Qinghao Hu 0001, Yifan Zhang 0001, Chunjie Zhang 0001, Yang Liu 0021, Jian Cheng 0001
CVPR6
2018 Block convolution: Towards memory-efficient inference of large-scale CNNs on FPGA
abstract
FPGA-based CNN accelerators are gaining popularity due to high energy efficiency and great flexibility in recent years. However, as the networks grow in depth and width, the great volume of intermediate data is too large to store on chip, data transfers between on-chip memory and off-chip memory should be frequently executed, which leads to unexpected offchip memory access latency and energy consumption. In this paper, we propose a block convolution approach, which is a memory-efficient, simple yet effective block-based convolution to completely avoid intermediate data from streaming out to off-chip memory during network inference. Experiments on the very large VGG-16 network show that the improved top-1/top-5 accuracy of 72.60%/91.10% can be achieved on the ImageNet classification task with the proposed approach. As a case study, we implement the VGG-16 network with block convolution on Xilinx Zynq ZC706 board, achieving a frame rate of 12.19fps under 150MHz working frequency, with all intermediate data staying on chip.
Gang Li 0015, Fanrong Li, Tianli Zhao, Jian Cheng 0001
DATE4
2018 Learning Compression from Limited Unlabeled Data
Jian Cheng 0001
ECCV (1)2
2018 Training Binary Weight Networks via Semi-Binary Decomposition
Qinghao Hu 0001, Gang Li 0015, Peisong Wang 0001, Yifan Zhang 0001, Jian Cheng 0001
ECCV (13)5
2018 Semi-supervised Generative Adversarial Hashing for Image Retrieval
Guan'an Wang, Qinghao Hu 0001, Jian Cheng 0001, Zeng-Guang Hou
ECCV (15)3
2018 BitStream: Efficient Computing Architecture for Real-Time Low-Power Inference of Binary Neural Networks on CPUs
abstract
Convolutional Neural Networks (CNN) have been widely used in many multimedia applications such as image classification, speech recognition, natural language processing and so on. Nevertheless, the high performance of deep CNN models also comes with high consumption of computation and storage resources, making it difficult to run CNN models in real time applications on mobile devices, where computational ability, memory resource and power are largely constrained. Binary network is a recently proposed tech- nique to reduce the computational and memory complexity of CNN, in which the expensive floating-point operations can be replaced by much cheaper bit-wise operations. Despite its obvious advantages, only few works explored the efficient implementation of Binary Neural Networks (BNN). In this work, we present a general architecture for efficient binary convolution referred to as BitStream with the latest computation flow for BNNs instead of the traditional row-major im2col based one. We mainly optimize memory access during computation of BNNs, the proposed calculation flow is cache friendly as well as can largely reduce memory overhead of BNNs, decidedly leading to its memory efficiency and further computational efficiency. Extensive evaluations on various networks demon- strate the efficiency of the proposed method. For instance, memory consumption of BitStream is reduced by 18-32× than original networks and 3× than existing implementations of BNNs during inference. Moreover, our implemented binary Alexnet can achieve 8.07× and 2.84× speedup over floating point precision model and conventional implementations of BNNs on 8 × Cortex A53 CPU s. With 4 × Intel CORE i7 6700 CPUs, the binary vgg-like convolutional network on CIFAR-10 runs even 1.69× faster than floating point precision version on TitanX GPU.
Tianli Zhao, Jian Cheng 0001
ACM Multimedia3
2018 Image-level classification by hierarchical structure learning with visual and semantic similarities
Chunjie Zhang 0001, Jian Cheng 0001, Qi Tian 0001
Inf. Sci.2
2018 Birds of a feather flock together: Visual representation with scale and class consistency
Chunjie Zhang 0001, Chenghua Li, Dongyuan Lu, Jian Cheng 0001, Qi Tian 0001
Inf. Sci.4
2018 Recent advances in efficient computation of deep convolutional neural networks
abstract
Deep neural networks have evolved remarkably over the past few years and they are currently the fundamental tools of many intelligent systems. At the same time, the computational complexity and resource consumption of these networks continue to increase. This poses a significant challenge to the deployment of such networks, especially in real-time applications or on resource-limited devices. Thus, network acceleration has become a hot topic within the deep learning community. As for hardware implementation of deep neural networks, a batch of accelerators based on a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC) have been proposed in recent years. In this paper, we provide a comprehensive survey of recent advances in network acceleration, compression, and accelerator design from both algorithm and hardware points of view. Specifically, we provide a thorough analysis of each of the following topics: network pruning, low-rank approximation, network quantization, teacher–student networks, compact network design, and hardware accelerators. Finally, we introduce and discuss a few possible future directions.
Jian Cheng 0001, Peisong Wang 0001, Gang Li 0015, Qinghao Hu 0001, Hanqing Lu
Frontiers Inf. Technol. Electron. Eng.1
2018 Incremental Codebook Adaptation for Visual Representation and Categorization
abstract
The bag-of-visual-words model is widely used for visual content analysis. For visual data, the codebook plays an important role for efficient representation. However, the codebook has to be relearned with the changes of training images. Once the codebook is changed, the encoding parameters of local features have to be recomputed. To alleviate this problem, in this paper, we propose an incremental codebook adaptation method for efficient visual representation. Instead of learning a new codebook, we gradually adapt a prelearned codebook using new images in an incremental way. To make use of the prelearned codebook, we try to make changes to the prelearned codebook with sparsity constraint and low-rank correlation. Besides, we also encode visually similar local features within a neighborhood to take advantage of locality information and ensure the encoded parameters are consistent. To evaluate the effectiveness of the proposed method, we apply the proposed method for categorization tasks on several public image datasets. Experimental results prove the effectiveness and usefulness of the proposed method over other codebook-based methods.
Chunjie Zhang 0001, Jian Cheng 0001, Qi Tian 0001
IEEE Trans. Cybern.2
2018 EgoGesture: A New Dataset and Benchmark for Egocentric Hand Gesture Recognition
abstract
Gesture is a natural interface in human-computer interaction, especially interacting with wearable devices, such as VR/AR helmet and glasses. However, in the gesture recognition community, it lacks of suitable datasets for developing egocentric (first-person view) gesture recognition methods, in particular in the deep learning era. In this paper, we introduce a new benchmark dataset named EgoGesture with sufficient size, variation, and reality to be able to train deep neural networks. This dataset contains more than 24 000 gesture samples and 3 000 000 frames for both color and depth modalities from 50 distinct subjects. We design 83 different static and dynamic gestures focused on interaction with wearable devices and collect them from six diverse indoor and outdoor scenes, respectively, with variation in background and illumination. We also consider the scenario when people perform gestures while they are walking. The performances of several representative approaches are systematically evaluated on two tasks: gesture classification in segmented data and gesture spotting and recognition in continuous data. Our empirical study also provides an in-depth analysis on input modality selection and domain adaptation between different scenes.
Yifan Zhang 0001, Congqi Cao, Jian Cheng 0001, Hanqing Lu
IEEE Trans. Multim.3
2018 Multiview Label Sharing for Visual Representations and Classifications
abstract
Different views represent different aspects of images. It is more effective to combine them for visual classifications. This paper proposes a novel multiview label sharing method to combine the discriminative power of different views for classifications. Especially, we linearly transfer different views into a shared space for representations. The inter-view similarities are kept in the shared space for each view. We also ensure the intra-view similarities of the same class between different views are preserved in the shared space. We jointly learn the classifiers and transformation matrices by minimizing the summed classification loss along with the inter-view and intra-view similarity constraints. In this paper, the inter-view constraints refer to the similarities between images of the corresponding view, whereas the intra-view constraints refer to the similarities between different views of images with the same semantics. Experimental results and analysis on several public datasets show the effectiveness of the proposed multiview label sharing method for visual classifications.
Chunjie Zhang 0001, Jian Cheng 0001, Qi Tian 0001
IEEE Trans. Multim.2
2018 Quantized CNN: A Unified Approach to Accelerate and Compress Convolutional Networks
abstract
We are witnessing an explosive development and widespread application of deep neural networks (DNNs) in various fields. However, DNN models, especially a convolutional neural network (CNN), usually involve massive parameters and are computationally expensive, making them extremely dependent on high-performance hardware. This prohibits their further extensions, e.g., applications on mobile devices. In this paper, we present a quantized CNN, a unified approach to accelerate and compress convolutional networks. Guided by minimizing the approximation error of individual layer's response, both fully connected and convolutional layers are carefully quantized. The inference computation can be effectively carried out on the quantized network, with much lower memory and storage consumption. Quantitative evaluation on two publicly available benchmarks demonstrates the promising performance of our approach: with comparable classification accuracy, it achieves 4 to $6 \times $ acceleration and 15 to $20\times $ compression. With our method, accurate image classification can even be directly carried out on mobile devices within 1 s.
Jian Cheng 0001, Jiaxiang Wu 0001, Cong Leng, Qinghao Hu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2018 Object Categorization Using Class-Specific Representations
abstract
Object categorization refers to the task of automatically classifying objects based on the visual content. Existing approaches simply represent each image with the visual features without considering the specific characters of images within the same class. However, objects of the same class may exhibit unique characters, which should be represented accordingly. In this brief, we propose a novel class-specific representation strategy for object categorization. For each class, we first model the characters of images within the same class using Gaussian mixture model (GMM). We then represent each image by calculating the Euclidean distance and relative Euclidean distance between the image and the GMM model for each class. We concatenate the representations of each class for joint representation. In this way, we can represent an image by not only considering the visual contents but also combining the class-specific characters. Experiments on several public available data sets validate the superiority of the proposed class-specific representation method over well-established algorithms for object category predictions.
Chunjie Zhang 0001, Jian Cheng 0001, Liang Li 0003, Qi Tian 0001
IEEE Trans. Neural Networks Learn. Syst.2
2018 Image-Specific Classification With Local and Global Discriminations
abstract
Most image classification methods try to learn classifiers for each class using training images alone. Due to the interclass and intraclass variations, it would be more effective to take the testing images into consideration for classifier learning. In this brief, we propose a novel image-specific classification method by combing the local and global discriminations of training images. We adaptively train classifier for each testing image instead of generating classifiers for each class with training images alone. For each testing image, we first select its ${k}$ nearest neighbors in the training set with the corresponding labels for local classifier training. This helps to model the distinctive characters of each testing image. Besides, we also use all the training images for global discrimination modeling. The local and global discriminations are combined for final classification. In this way, we could not only model the specific character of each testing image but also avoid the local optimum by jointly considering all the training images. To evaluate the usefulness of the proposed image-specific classification with local and global discrimination (ISC-LG) method, we conduct image classification experiments on several public image data sets. The superior performances over other baseline methods prove the effectiveness of the proposed ISC-LG method.
Chunjie Zhang 0001, Jian Cheng 0001, Qi Tian 0001
IEEE Trans. Neural Networks Learn. Syst.2
2018 Structured Weak Semantic Space Construction for Visual Categorization
abstract
Visual features have been widely used for image representation and categorization. However, visual features are often inconsistent with human perception. Besides, constructing explicit semantic space is still an open problem. To alleviate these two problems, in this paper, we propose to construct structured weak semantic space for image representation. Exemplar classifier is first trained to separate each training image from other images for weak semantic space construction. However, each exemplar classifier separates one training image from other images, and it only has limited semantic separability. Besides, the outputs of exemplar classifiers are inconsistent with each other. We jointly construct the weak semantic space using structured constraint. This is achieved by imposing low-rank constraint on the outputs of exemplar classifiers with sparsity constraint. An alternative optimization procedure is used to learn the exemplar classifiers. Since the proposed method does not dependent on the initial image representation strategy, we can make use of various visual features for efficient exemplar classifier training (e.g., fisher vector-based methods and convolutional neural networks-based methods). We apply the proposed structured weak semantic space-based image representation method for categorization. The experimental results on several public image data sets prove the effectiveness of the proposed method.
Chunjie Zhang 0001, Jian Cheng 0001, Qi Tian 0001
IEEE Trans. Neural Networks Learn. Syst.2
2018 DeepSearch: A Fast Image Search Framework for Mobile Devices
abstract
Content-based image retrieval (CBIR) is one of the most important applications of computer vision. In recent years, there have been many important advances in the development of CBIR systems, especially Convolutional Neural Networks (CNNs) and other deep-learning techniques. On the other hand, current CNN-based CBIR systems suffer from high computational complexity of CNNs. This problem becomes more severe as mobile applications become more and more popular. The current practice is to deploy the entire CBIR systems on the server side while the client side only serves as an image provider. This architecture can increase the computational burden on the server side, which needs to process thousands of requests per second. Moreover, sending images have the potential of personal information leakage. As the need of mobile search expands, concerns about privacy are growing. In this article, we propose a fast image search framework, named DeepSearch, which makes complex image search based on CNNs feasible on mobile phones. To implement the huge computation of CNN models, we present a tensor Block Term Decomposition (BTD) approach as well as a nonlinear response reconstruction method to accelerate the CNNs involving in object detection and feature extraction. The extensive experiments on the ImageNet dataset and Alibaba Large-scale Image Search Challenge dataset show that the proposed accelerating approach BTD can significantly speed up the CNN models and further makes CNN-based image search practical on common smart phones.
Peisong Wang 0001, Qinghao Hu 0001, Zhiwei Fang, Chaoyang Zhao, Jian Cheng 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2017 Fast K-means for Large Scale Clustering
abstract
K-means algorithm has been widely used in machine learning and data mining due to its simplicity and good performance. However, the standard k-means algorithm would be quite slow for clustering millions of data into thousands of or even tens of thousands of clusters. In this paper, we propose a fast k-means algorithm named multi-stage k-means (MKM) which uses a multi-stage filtering approach. The multi-stage filtering approach greatly accelerates the k-means algorithm via a coarse-to-fine search strategy. To further speed up the algorithm, hashing is introduced to accelerate the assignment step which is the most time-consuming part in k-means. Extensive experiments on several massive datasets show that the proposed algorithm can obtain up to 600X speed-up over the k-means algorithm with comparable accuracy.
Qinghao Hu 0001, Jiaxiang Wu 0001, Lu Bai 0001, Yifan Zhang 0001, Jian Cheng 0001
CIKM5
2017 Fixed-Point Factorized Networks
abstract
In recent years, Deep Neural Networks (DNN) based methods have achieved remarkable performance in a wide range of tasks and have been among the most powerful and widely used techniques in computer vision. However, DNN-based methods are both computational-intensive and resource-consuming, which hinders the application of these methods on embedded systems like smart phones. To alleviate this problem, we introduce a novel Fixed-point Factorized Networks (FFN) for pretrained models to reduce the computational complexity as well as the storage requirement of networks. The resulting networks have only weights of -1, 0 and 1, which significantly eliminates the most resource-consuming multiply-accumulate operations (MACs). Extensive experiments on large-scale ImageNet classification task show the proposed FFN only requires one-thousandth of multiply operations with comparable accuracy.
Peisong Wang 0001, Jian Cheng 0001
CVPR2
2017 Egocentric Gesture Recognition Using Recurrent 3D Convolutional Neural Networks with Spatiotemporal Transformer Modules
abstract
Gesture is a natural interface in interacting with wearable devices such as VR/AR helmet and glasses. The main challenge of gesture recognition in egocentric vision arises from the global camera motion caused by the spontaneous head movement of the device wearer. In this paper, we address the problem by a novel recurrent 3D convolutional neural network for end-to-end learning. We specially design a spatiotemporal transformer module with recurrent connections between neighboring time slices which can actively transform a 3D feature map into a canonical view in both spatial and temporal dimensions. To validate our method, we introduce a new dataset with sufficient size, variation and reality, which contains 83 gestures designed for interaction with wearable devices, and more than 24,000 RGB-D gesture samples from 50 subjects captured in 6 scenes. On this dataset, we show that the proposed network outperforms competing state-of-the-art algorithms. Moreover, our method can achieve state-of-the-art performance on the challenging GTEA egocentric action dataset.
Congqi Cao, Yifan Zhang 0001, Yi Wu 0001, Hanqing Lu, Jian Cheng 0001
ICCV5
2017 De-biased dart ensemble model for personalized recommendation
abstract
Personalized recommendation aims to use the historical behavior of users to recommend new items that are likely to be of interest to them. Due to a tiny improvement of it can lead to a huge profits, lots of giant e-commerce companies, such as Amazon, Alibaba and eBay, have put their great effort on this field. In this paper, we formulate the problem of personalized recommendation as a tree based regression task and present a de-biased DART ensemble model. To generate compact personalized profile of user, a decision tree based sparse encoding approach is first proposed to refine those raw behaviors, like clicks, views, and purchases from user-item interaction. Considering the problem of class imbalance in training the predictive models, resampling based de-biased DART model is utilized to form multiple weak predictors for recommendation. Thus, based on these learnt weak predictors, a stronger predictor can be built. Experimental results on real commercial dataset demonstrate the effectiveness of the proposed personalized recommendation model.
Deqiang Kong, Jingyuan Tang, Zhenfeng Zhu, Jian Cheng 0001, Yao Zhao 0001
ICME4
2017 Measuring Word Semantic Similarity Based on Transferred Vectors
Changliang Li, Yujun Zhou 0001, Jian Cheng 0001, Bo Xu 0002
ICONIP (4)4
2017 Pseudo Label based Unsupervised Deep Discriminative Hashing for Image Retrieval
abstract
Hashing methods play an important role in large scale image retrieval. Traditional hashing methods use hand-crafted features to learn hash functions, which can not capture the high level semantic information. Deep hashing algorithms use deep neural networks to learn feature representation and hash functions simultaneously. Most of these algorithms exploit supervised information to train the deep network. However, supervised information is expensive to obtain. In this paper, we propose a pseudo label based unsupervised deep discriminative hashing algorithm. First, we cluster images via K-means and the cluster labels are treated as pseudo labels. Then we train a deep hashing network with pseudo labels by minimizing the classification loss and quantization loss. Experiments on two datasets demonstrate that our unsupervised deep discriminative hashing method outperforms the state-of-art unsupervised hashing methods.
Qinghao Hu 0001, Jiaxiang Wu 0001, Jian Cheng 0001, Lifang Wu, Hanqing Lu
ACM Multimedia3
2016 Shoot to Know What: An Application of Deep Networks on Mobile Devices
abstract
Convolutional neural networks (CNNs) have achieved impressive performance in a wide range of computer vision areas. However, the application on mobile devices remains intractable due to the high computation complexity. In this demo, we propose the Quantized CNN (Q-CNN), an efficient framework for CNN models, to fulfill efficient and accurate image classification on mobile devices. Our Q-CNN framework dramatically accelerates the computation and reduces the storage/memory consumption, so that mobile devices can independently run an ImageNet-scale CNN model. Experiments on the ILSVRC-12 dataset demonstrate 4~6x speed-up and 15~20x compression, with merely one percentage drop in the classification accuracy. Based on the Q-CNN framework, even mobile devices can accurately classify images within one second.
Jiaxiang Wu 0001, Qinghao Hu 0001, Cong Leng, Jian Cheng 0001
AAAI4
2016 Quantized Convolutional Neural Networks for Mobile Devices
abstract
Recently, convolutional neural networks (CNN) have demonstrated impressive performance in various computer vision tasks. However, high performance hardware is typically indispensable for the application of CNN models due to the high computation complexity, which prohibits their further extensions. In this paper, we propose an efficient framework, namely Quantized CNN, to simultaneously speed-up the computation and reduce the storage and memory overhead of CNN models. Both filter kernels in convolutional layers and weighting matrices in fully-connected layers are quantized, aiming at minimizing the estimation error of each layer's response. Extensive experiments on the ILSVRC-12 benchmark demonstrate 4 ~ 6× speed-up and 15 ~ 20× compression with merely one percentage loss of classification accuracy. With our quantized CNN model, even mobile devices can accurately classify images within one second.
Jiaxiang Wu 0001, Cong Leng, Qinghao Hu 0001, Jian Cheng 0001
CVPR5
2016 Accelerating Convolutional Neural Networks for Mobile Applications
abstract
Convolutional neural networks (CNNs) have achieved remarkable performance in a wide range of computer vision tasks, typically at the cost of massive computational complexity. The low speed of these networks may hinder real-time applications especially when computational resources are limited. In this paper, an efficient and effective approach is proposed to accelerate the test-phase computation of CNNs based on low-rank and group sparse tensor decomposition. Specifically, for each convolutional layer, the kernel tensor is decomposed into the sum of a small number of low multilinear rank tensors. Then we replace the original kernel tensors in all layers with the approximate tensors and fine-tune the whole net with respect to the final classification task using standard backpropagation. \\ Comprehensive experiments on ILSVRC-12 demonstrate significant reduction in computational complexity, at the cost of negligible loss in accuracy. For the widely used VGG-16 model, our approach obtains a 6.6$\times$ speed-up on PC and 5.91$\times$ speed-up on mobile device of the whole network with less than 1\% increase on top-5 error.
Peisong Wang 0001, Jian Cheng 0001
ACM Multimedia2
2016 LSSLP - Local structure sensitive label propagation
Zhenfeng Zhu, Jian Cheng 0001, Yao Zhao 0001, Jieping Ye
Inf. Sci.2
2016 Enriching one-class collaborative filtering with content information from social media
Jian Cheng 0001, Xi Zhang 0018, Qinshan Liu, Hanqing Lu
Multim. Syst.2
2016 Guest Editorial: Image Analysis and Processing Leveraging Additional Information
Luis Herranz, Jian Cheng 0001, Yue Gao 0002, Shuqiang Jiang
Multim. Tools Appl.2
2015 Personalized Recommendation Meets Your Next Favorite
abstract
A comprehensive understanding of user's item selection behavior is not only essential to many scientific disciplines, but also has a profound business impact on online recommendation. Recent researches have discovered that user's favorites can be divided into 2 categories: long-term and short-term. User's item selection behavior is a mixed decision of her long and short-term favorites. In this paper, we propose a unified model, namely States Transition pAir-wise Ranking Model (STAR), to address users' favorites mining for sequential-set recommendation. Our method utilizes a transition graph for collaborative filtering that accounts for mining user's short-term favorites, jointed with a generative topic model for expressing user's long-term favorites. Furthermore, a user's specific prior is introduced into our unified model for better modeling personalization. Technically, we develop a pair-wise ranking loss function for parameters learning. Empirically, we measure the effectiveness of our method using two real-world datasets and the results show that our method outperforms state-of-the-art methods.
Jian Cheng 0001, Hanqing Lu
CIKM2
2015 Online sketching hashing
abstract
Recently, hashing based approximate nearest neighbor (ANN) search has attracted much attention. Extensive new algorithms have been developed and successfully applied to different applications. However, two critical problems are rarely mentioned. First, in real-world applications, the data often comes in a streaming fashion but most of existing hashing methods are batch based models. Second, when the dataset becomes huge, it is almost impossible to load all the data into memory to train hashing models. In this paper, we propose a novel approach to handle these two problems simultaneously based on the idea of data sketching. A sketch of one dataset preserves its major characters but with significantly smaller size. With a small size sketch, our method can learn hash functions in an online fashion, while needs rather low computational complexity and storage space. Extensive experiments on two large scale benchmarks and one synthetic dataset demonstrate the efficacy of the proposed method.
Cong Leng, Jiaxiang Wu 0001, Jian Cheng 0001, Xiao Bai 0001, Hanqing Lu
CVPR3
2015 Hashing for Distributed Data
abstract
Recently, hashing based approximate nearest neighbors search has attracted much attention. Extensive centralized hashing algorithms have been proposed and achieved promising performance. However, due to the large scale of many applications, the data is often stored or even collected in a distributed manner. Learning hash functions by aggregating all the data into a fusion center is infeasible because of the prohibitively expensive communication and computation overhead. In this paper, we develop a novel hashing model to learn hash functions in a distributed setting. We cast a centralized hashing model as a set of subproblems with consensus constraints. We find these subproblems can be analytically solved in parallel on the distributed compute nodes. Since no training data is transmitted across the nodes in the learning process, the communication cost of our model is independent to the data size. Extensive experiments on several large scale datasets containing up to 100 million samples demonstrate the efficacy of our method.
Cong Leng, Jiaxiang Wu 0001, Jian Cheng 0001, Xi Zhang 0018, Hanqing Lu
ICML3
2015 Learning Deep Features For MSR-bing Information Retrieval Challenge
abstract
Two tasks have been put forward in the MSR-bing Grand Challenge 2015. To address the information retrieval task, we raise and integrate a series of methods with visual features obtained by convolution neural network (CNN) models. In our experiments, we discover that the ranking strategies of Hierarchical clustering and PageRank methods are mutually complementary. Another task is fine-grained classification. In contrast to basic-level recognition, fine-grained classification aims to distinguish between different breeds or species or product models, and often requires distinctions that must be conditioned on the object pose for reliable identification. Current state-of-the-art techniques rely heavily upon the use of part annotations, while the bing datasets suffer both abundance of part annotations and dirty background. In this paper, we propose a CNN-based feature representation for visual recognition only using image-level information. Our CNN model is pre-trained on a collection of clean datasets and fine-tuned on the bing datasets. Furthermore, a multi-scale training strategy is adopted by simply resizing the input images into different scales and then merging the soft-max posteriors. We then implement our method into a unified visual recognition system on Microsoft cloud service. Finally, our solution achieved top performance in both tasks of the contest
Sixie Yu, Cong Leng, Jiaxiang Wu 0001, Qinghao Hu 0001, Jian Cheng 0001
ACM Multimedia6
2015 Incremental Matrix Factorization via Feature Space Re-learning for Recommender System
abstract
Matrix factorization is widely used in Recommender Systems. Although existing popular incremental matrix factorization methods are effectively in reducing time complexity, they simply assume that the similarity between items or users is invariant. For instance, they keep the item feature matrix unchanged and just update the user matrix without re-training the entire model. However, with the new users growing continuously, the fitting error would be accumulated since the extra distribution information of items has not been utilized. In this paper, we present an alternative and reasonable approach, with a relaxed assumption that the similarity between items (users) is relatively stable after updating. Concretely, utilizing the prediction error of the new data as the auxiliary features, our method updates both feature matrices simultaneously, and thus users' preference can be better modeled than merely adjusting one corresponded feature matrix. Besides, our method maintains the feature dimension in a smaller size through taking advantage of matrix sketching. Experimental results show that our proposal outperforms the existing incremental matrix factorization methods.
Jian Cheng 0001, Hanqing Lu
RecSys2
2015 When Personalization Meets Conformity: Collective Similarity based Multi-Domain Recommendation
abstract
Existing recommender systems place emphasis on personalization to achieve promising accuracy. However, in the context of multiple domain, users are likely to seek the same behaviors as domain authorities. This conformity effect provides a wealth of prior knowledge when it comes to multi-domain recommendation, but has not been fully exploited. In particular, users whose behaviors are significant similar with the public tastes can be viewed as domain authorities. To detect these users meanwhile embed conformity into recommendation, a domain-specific similarity matrix is intuitively employed. Therefore, a collective similarity is obtained to leverage the conformity with personalization. In this paper, we establish a Collective Structure Sparse Representation(CSSR) method for multi-domain recommendation. Based on adaptive $k$-Nearest-Neighbor framework, we impose the lasso and group lasso penalties as well as least square loss to jointly optimize the collective similarity. Experimental results on real-world data confirm the effectiveness of the proposed method.
Xi Zhang 0018, Jian Cheng 0001, Shuang Qiu 0002, Zhenfeng Zhu, Hanqing Lu
SIGIR2
2015 Image classification using boosted local features with random orientation and location selection
Chunjie Zhang 0001, Jian Cheng 0001, Yifan Zhang 0001, Jing Liu 0001, Chao Liang 0001, Junbiao Pang, Qingming Huang, Qi Tian 0001
Inf. Sci.2
2015 How friends affect user behaviors? An exploration of social relation analysis for recommendation
Jian Cheng 0001, Xi Zhang 0018, Qingshan Liu 0001, Hanqing Lu
Knowl. Based Syst.2
2015 DualDS: A dual discriminative rating elicitation framework for cold start recommendation
Xi Zhang 0018, Jian Cheng 0001, Shuang Qiu 0002, Guibo Zhu, Hanqing Lu
Knowl. Based Syst.2
2015 Consensus hashing
Cong Leng, Jian Cheng 0001
Mach. Learn.2
2015 Learning latent semantic model with visual consistency for image analysis
Jian Cheng 0001, Ting Rui, Hanqing Lu
Multim. Tools Appl.1
2015 Beyond Explicit Codebook Generation: Visual Representation Using Implicitly Transferred Codebooks
abstract
The bag-of-visual-words model plays a very important role for visual applications. Local features are first extracted and then encoded to get the histogram-based image representation. To encode local features, a proper codebook is needed. Usually, the codebook has to be generated for each data set which means the codebook is data set dependent. Besides, the codebook may be biased when we only have a limited number of training images. Moreover, the codebook has to be pre-learned which cannot be updated quickly, especially when applied for online visual applications. To solve the problems mentioned above, in this paper, we propose a novel implicit codebook transfer method for visual representation. Instead of explicitly generating the codebook for the new data set, we try to make use of pre-learned codebooks using non-linear transfer. This is achieved by transferring the pre-learned codebooks with non-linear transformation and use them to reconstruct local features with sparsity constraints. The codebook does not need to be explicitly generated but can be implicitly transferred. In this way, we are able to make use of pre-learned codebooks for new visual applications by implicitly learning the codebook and the corresponding encoding parameters for image representation. We apply the proposed method for image classification and evaluate the performance on several public image data sets. Experimental results demonstrate the effectiveness and efficiency of the proposed method.
Chunjie Zhang 0001, Jian Cheng 0001, Jing Liu 0001, Junbiao Pang, Qingming Huang, Qi Tian 0001
IEEE Trans. Image Process.2
2014 Recommendation by Mining Multiple User Behaviors with Group Sparsity
abstract
Recently, some recommendation methods try to improvethe prediction results by integrating informationfrom user’s multiple types of behaviors. How to modelthe dependence and independence between differentbehaviors is critical for them. In this paper, we proposea novel recommendation model, the Group-Sparse MatrixFactorization (GSMF), which factorizes the ratingmatrices for multiple behaviors into the user and itemlatent factor space with group sparsity regularization.It can (1) select out the different subsets of latent factorsfor different behaviors, addressing that users’ decisionson different behaviors are determined by differentsets of factors;(2) model the dependence and independencebetween behaviors by learning the sharedand private factors for multiple behaviors automatically; (3) allow the shared factors between different behaviorsto be different, instead of all the behaviors sharingthe same set of factors. Experiments on the real-world dataset demonstrate that our model can integrate users’multiple types of behaviors into recommendation better,compared with other state-of-the-arts.
Jian Cheng 0001, Xi Zhang 0018, Shuang Qiu 0002, Hanqing Lu
AAAI2
2014 Supervised Hashing with Soft Constraints
abstract
Due to the ability to preserve semantic similarity in Hamming space, supervised hashing has been extensively studied recently. Most existing approaches encourage two dissimilar samples to have maximum Hamming distance. This may lead to an unexpected consequence that two unnecessarily similar samples would have the same code if they are both dissimilar with another sample. Besides, in existing methods, all labeled pairs are treated with equal importance without considering the semantic gap, which is not conducive to thoroughly leverage the supervised information. We present a general framework for supervised hashing to address the above two limitations. We do not toughly require a dissimilar pair to have maximum Hamming distance. Instead, a soft constraint which can be viewed as a regularization to avoid over-fitting is utilized. Moreover, we impose different weights to different training pairs, and these weights can be automatically adjusted in the learning process. Experiments on two benchmarks show that the proposed method can easily outperform other state-of-the-art methods.
Cong Leng, Jian Cheng 0001, Jiaxiang Wu 0001, Xi Zhang 0018, Hanqing Lu
CIKM2
2014 Fast and Accurate Image Matching with Cascade Hashing for 3D Reconstruction
abstract
Image matching is one of the most challenging stages in 3D reconstruction, which usually occupies half of computational cost and inaccurate matching may lead to failure of reconstruction. Therefore, fast and accurate image matching is very crucial for 3D reconstruction. In this paper, we proposed a Cascade Hashing strategy to speed up the image matching. In order to accelerate the image matching, the proposed Cascade Hashing method is designed to be three-layer structure: hashing lookup, hashing remapping, and hashing ranking. Each layer adopts different measures and filtering strategies, which is demonstrated to be less sensitive to noise. Extensive experiments show that image matching can be accelerated by our approach in hundreds times than brute force matching, even achieves ten times or more than Kd-tree based matching while retaining comparable accuracy.
Jian Cheng 0001, Cong Leng, Jiaxiang Wu 0001, Hainan Cui, Hanqing Lu
CVPR1
2014 Adaptive Object Retrieval with Kernel Reconstructive Hashing
abstract
Hashing is very useful for fast approximate similarity search on large database. In the unsupervised settings, most hashing methods aim at preserving the similarity defined by Euclidean distance. Hash codes generated by these approaches only keep their Hamming distance corresponding to the pairwise Euclidean distance, ignoring the local distribution of each data point. This objective does not hold for k-nearest neighbors search. In this paper, we firstly propose a new adaptive similarity measure which is consistent with k-NN search, and prove that it leads to a valid kernel. Then we propose a hashing scheme which uses binary codes to preserve the kernel function. Using low-rank approximation, our hashing framework is more effective than existing methods that preserve similarity over arbitrary kernel. The proposed kernel function, hashing framework, and their combination have demonstrated significant advantages compared with several state-of-the-art methods.
Haichuan Yang, Xiao Bai 0001, Jun Zhou 0001, Peng Ren 0001, Zhihong Zhang 0001, Jian Cheng 0001
CVPR6
2014 Semi-randomized hashing for large scale data retrieval
abstract
In information retrieval, efficient accomplishing the nearest neighbor search on large scale database is a great challenge. Hashing based indexing methods represent each data instance as a binary string to retrieve the approximate nearest neighbors. In this paper, we present a semi-randomized hashing approach to preserve the Euclidean distance by binary codes. Euclidean distance preserving is a classic research problem in hashing. Most hashing methods used purely randomized or optimized learning strategy to achieve this goal. Our method, on the other hand, combines both randomized and optimized strategies. It starts from generating multiple random vectors, and then approximates them by a single projection vector. In the quantization step, it uses the orthogonal transformation to minimize an upper bound of the deviation between real-valued vectors and binary codes. The proposed method overcomes the problem that randomized hash functions are isolated from the data distribution. What's more, our method supports an arbitrary number of hash functions, which is beneficial in building better hashing methods. The experiments show that our approach outperforms the alternative state-of-the-art methods for retrieval on the large scale dataset.
Haichuan Yang, Xiao Bai 0001, Jun Zhou 0001, Peng Ren 0001, Jian Cheng 0001, Lu Bai 0001
DSAA5
2014 Community discovering guided cold-start recommendation: A discriminative approach
abstract
Recommendation for new users is a key challenge due to the lack of prior information from them, which is the well-known cold-start problem. Preference elicitation has been proposed as an efficient strategy for eliciting new users preference through an initial interview where new users are queried by elaborately selected items. In this paper, we propose a novel community discovering guided discriminative selection (CDDS) model for constructing query set. We exploit the community as an effective information which is not fully used in existing approaches. By integrating item selection and community discovery into one framework, our model selects most discriminative items for preference elicitation, with guidance of unsupervised community discovering process. To perform community discovering process, the model utilizes rating similarity graph and social network as a graph regular-ization. Experimental results on real-world datasets Flixster and Douban demonstrate that the proposed method outperforms traditional preference elicitation methods for cold-start recommendation.
Shuang Qiu 0002, Jian Cheng 0001, Xi Zhang 0018, Biao Niu, Hanqing Lu
ICME2
2014 Learning Binary Codes with Bagging PCA
Cong Leng, Jian Cheng 0001, Xiao Bai 0001, Hanqing Lu
ECML/PKDD (2)2
2014 Group latent factor model for recommendation with multiple user behaviors
abstract
Recently, some recommendation methods try to relieve the data sparsity problem of Collaborative Filtering by exploiting data from users' multiple types of behaviors. However, most of the exist methods mainly consider to model the correlation between different behaviors and ignore the heterogeneity of them, which may make improper information transferred and harm the recommendation results. To address this problem, we propose a novel recommendation model, named Group Latent Factor Model (GLFM), which attempts to learn a factorization of latent factor space into subspaces that are shared across multiple behaviors and subspaces that are specific to each type of behaviors. Thus, the correlation and heterogeneity of multiple behaviors can be modeled by these shared and specific latent factors. Experiments on the real-world dataset demonstrate that our model can integrate users' multiple types of behaviors into recommendation better.
Jian Cheng 0001, Jinqiao Wang, Hanqing Lu
SIGIR1
2014 Random subspace for binary codes learning in large scale image retrieval
abstract
Due to the fast query speed and low storage cost, hashing based approximate nearest neighbor search methods have attracted much attention recently. Many state of the art methods are based on eigenvalue decomposition. In these approaches, the information caught in different dimensions is unbalanced and generally most of the information is contained in the top eigenvectors. We demonstrate that this leads to an unexpected phenomenon that longer hashing code does not necessarily yield better performance. In this work, we introduce a random subspace strategy to address this limitation. At first, a small fraction of the whole feature space is randomly sampled to train the hashing algorithms each time and only the top eigenvectors are kept to generate one piece of short code. This process will be repeated several times and then the obtained many pieces of short codes are concatenated into one piece of long code. Theoretical analysis and experiments on two benchmarks confirm the effectiveness of the proposed strategy for hashing.
Cong Leng, Jian Cheng 0001, Hanqing Lu
SIGIR2
2014 Item group based pairwise preference learning for personalized ranking
abstract
Collaborative filtering with implicit feedbacks has been steadily receiving more attention, since the abundant implicit feedbacks are more easily collected while explicit feedbacks are not necessarily always available. Several recent work address this problem well utilizing pairwise ranking method with a fundamental assumption that a user prefers items with positive feedbacks to the items without observed feedbacks, which also implies that the items without observed feedbacks are treated equally without distinction. However, users have their own preference on different items with different degrees which can be modeled into a ranking relationship. In this paper, we exploit this prior information of a user's preference from the nearest neighbor set by the neighbors' implicit feedbacks, which can split items into different item groups with specific ranking relations. We propose a novel PRIGP(Personalized Ranking with Item Group based Pairwise preference learning) algorithm to integrate item based pairwise preference and item group based pairwise preference into the same framework. Experimental results on three real-world datasets demonstrate the proposed method outperforms the competitive baselines on several ranking-oriented evaluation metrics.
Shuang Qiu 0002, Jian Cheng 0001, Cong Leng, Hanqing Lu
SIGIR2
2014 Semi-supervised multi-graph hashing for scalable similarity search
Jian Cheng 0001, Cong Leng, Meng Wang 0001, Hanqing Lu
Comput. Vis. Image Underst.1
2014 Beyond semantic attributes: Auxiliary feature discovery for image classification
Biao Niu, Jian Cheng 0001, Yang Liu 0021, Hanqing Lu
Neurocomputing2
2014 Object categorization in sub-semantic space
Chunjie Zhang 0001, Jian Cheng 0001, Jing Liu 0001, Junbiao Pang, Chao Liang 0001, Qingming Huang, Qi Tian 0001
Neurocomputing2
2014 Bayesian co-boosting for multi-modal gesture recognition
Jiaxiang Wu 0001, Jian Cheng 0001
J. Mach. Learn. Res.2
2014 Data-Dependent Hashing Based on p-Stable Distribution
abstract
The p-stable distribution is traditionally used for data-independent hashing. In this paper, we describe how to perform data-dependent hashing based on p-stable distribution. We commence by formulating the Euclidean distance preserving property in terms of variance estimation. Based on this property, we develop a projection method, which maps the original data to arbitrary dimensional vectors. Each projection vector is a linear combination of multiple random vectors subject to p-stable distribution, in which the weights for the linear combination are learned based on the training data. An orthogonal matrix is then learned data-dependently for minimizing the thresholding error in quantization. Combining the projection method and orthogonal matrix, we develop an unsupervised hashing scheme, which preserves the Euclidean distance. Compared with data-independent hashing methods, our method takes the data distribution into consideration and gives more accurate hashing results with compact hash codes. Different from many data-dependent hashing methods, our method accommodates multiple hash tables and is not restricted by the number of hash functions. To extend our method to a supervised scenario, we incorporate a supervised label propagation scheme into the proposed projection method. This results in a supervised hashing scheme, which preserves semantic similarity of data. Experimental results show that our methods have outperformed several state-of-the-art hashing approaches in both effectiveness and efficiency.
Xiao Bai 0001, Haichuan Yang, Jun Zhou 0001, Peng Ren 0001, Jian Cheng 0001
IEEE Trans. Image Process.5
2013 Semi-supervised discriminative preference elicitation for cold-start recommendation
abstract
Recommendation for cold users is fairly challenging because no prior rating can be used in preference prediction. To tackle this cold-start scenario, rating elicitation is usually employed through an initial interview in which users are queried by some carefully selected items. In this paper, we propose a novel framework to mine the most valuable items to construct query set using a semi-supervised discriminative selection (SSDS) model. To learn a low dimensional representation for users in item space which can reflect their tastes to a large extent, the model incorporates category labels as discriminative information. To ensure the used labels reliable as well as all users considered, the model utilizes a semi-supervised scheme leveraging expert guidance with graph regularization. Experimental results on real-world dataset MovieLens demonstrate that the proposed SSDS model outperforms traditional preference elicitation methods on top-N measures for cold-start recommendation.
Xi Zhang 0018, Jian Cheng 0001, Biao Niu, Hanqing Lu
CIKM2
2013 Cost-sensitive background subtraction
abstract
Foreground and background are treated without distinction at classification stage in most background subtraction algorithms. However, correct classification of foreground is the primary requirement, and thus misclassification costs of the two classes should be different. Based on this fact, we present a new method to introduce cost sensitivity into background subtraction, where a cost matrix is created to represent the costs of misclassification. Some items in the cost matrix are not constants, but functions of foreground occurence at each pixel location. By the use of such non-constant costs, detection rate of foreground is improved while increase of false alarms is prevented at the same time. Experiments demonstrate the effectiveness of the proposed algorithm.
Xiang Zhang 0006, Jian Cheng 0001, Zhi Liu 0003, Jie Yang 0002
ICIP2
2013 Fusing multi-modal features for gesture recognition
abstract
This paper proposes a novel multi-modal gesture recognition framework and introduces its application to continuous sign language recognition. A Hidden Markov Model is used to construct the audio feature classifier. A skeleton feature classifier is trained to provided complementary information based on the Dynamic Time Warping model. The confidence scores generated by two classifiers are firstly normalized and then combined to produce a weighted sum for the final recognition. Experimental results have shown that the precision and recall scores for 20 classes of our multi-modal recognition framework can achieve 0.8829 and 0.8890 respectively, which proves that our method is able to correctly reject false detection caused by single classifier. Our approach scored 0.12756 in mean Levenshtein distance and was ranked 1st in the Multi-modal Gesture Recognition Challenge in 2013.
Jiaxiang Wu 0001, Jian Cheng 0001, Chaoyang Zhao, Hanqing Lu
ICMI2
2013 A Weighted One Class Collaborative Filtering with Content Topic Features
Jian Cheng 0001, Xi Zhang 0018, Qingshan Liu 0001, Hanqing Lu
MMM (2)2
2013 TopRec: domain-specific recommendation through community topic mining in social network
abstract
Traditionally, Collaborative Filtering assumes that similar users have similar responses to similar items. However, human activities exhibit heterogenous features across multiple domains such that users own similar tastes in one domain may behave quite differently in other domains. Moreover, highly sparse data presents crucial challenge in preference prediction. Intuitively, if users' interested domains are captured first, the recommender system is more likely to provide the enjoyed items while filter out those uninterested ones. Therefore, it is necessary to learn preference profiles from the correlated domains instead of the entire user-item matrix. In this paper, we propose a unified framework, TopRec, which detects topical communities to construct interpretable domains for domain-specific collaborative filtering. In order to mine communities as well as the corresponding topics, a semi-supervised probabilistic topic model is utilized by integrating user guidance with social network. Experimental results on real-world data from Epinions and Ciao demonstrate the effectiveness of the proposed framework.
Xi Zhang 0018, Jian Cheng 0001, Biao Niu, Hanqing Lu
WWW2
2013 Hashing with dual complementary projection learning for fast image retrieval
Jian Cheng 0001, Hanqing Lu
Neurocomputing2
2013 Hierarchical Remote Sensing Image Analysis via Graph Laplacian Energy
abstract
Segmentation and classification are important tasks in remote sensing image analysis. Recent research shows that images can be described in hierarchical structure or regions. Such hierarchies can produce the state-of-the-art segmentations and can be used in the classification. However, they often contain more levels and regions than required for an efficient image description, which may cause increased computational complexity. In this letter, we propose a new hierarchical segmentation method that applies graph Laplacian energy as a generic measure for segmentation. It reduces the redundancy in the hierarchy by an order of magnitude with little or no loss of performance. In the classification stage, we apply local self-similarity feature to capture the internal geometric layouts of regions in an image. By incorporating advantages from both semantic hierarchical segmentation and local geometric region description, we have achieved better performance than those from the methods being compared. In the experimental section, we validate the effectiveness of our method by showing results on QuickBird and GeoEye-1 image data sets.
Huigang Zhang, Xiao Bai 0001, Huaxin Zheng, Huijie Zhao, Jun Zhou 0001, Jian Cheng 0001, Hanqing Lu
IEEE Geosci. Remote. Sens. Lett.6
2013 Dual local consistency hashing with discriminative projections selection
Jian Cheng 0001, Hanqing Lu
Signal Process.2
2013 Asymmetric propagation based batch mode active learning for image retrieval
Biao Niu, Jian Cheng 0001, Xiao Bai 0001, Hanqing Lu
Signal Process.2
2013 Object Detection Via Structural Feature Selection and Shape Model
abstract
In this paper, we propose an approach for object detection via structural feature selection and part-based shape model. It automatically learns a shape model from cluttered training images without need to explicitly use bounding boxes on objects. Our approach first builds a class-specific codebook of local contour features, and then generates structural feature descriptors by combining context shape information. These descriptors are robust to both within-class variations and scale changes. Through exploring pairwise image matching using fast earth mover's distance, feature weights can be iteratively updated. Those discriminative foreground features are assigned high weights and then selected to build a part-based shape model. Finally, object detection is performed by matching each testing image with this model. Experiments show that the proposed method is very effective. It has achieved comparable performance to the state-of-the-art shape-based detection methods, but requires much less training information.
Huigang Zhang, Xiao Bai 0001, Jun Zhou 0001, Jian Cheng 0001, Huijie Zhao
IEEE Trans. Image Process.4
2013 Spectral Hashing With Semantically Consistent Graph for Image Indexing
abstract
The ability of fast similarity search in a large-scale dataset is of great importance to many multimedia applications. Semantic hashing is a promising way to accelerate similarity search, which designs compact binary codes for a large number of images so that semantically similar images are mapped to close codes. Retrieving similar neighbors is then simply accomplished by retrieving images that have codes within a small Hamming distance of the code of the query. Among various hashing approaches, spectral hashing (SH) has shown promising performance by learning the binary codes with a spectral graph partitioning method. However, the Euclidean distance is usually used to construct the graph Laplacian in SH, which may not reflect the inherent distribution of the data. Therefore, in this paper, we propose a method to directly optimize the graph Laplacian. The learned graph, which can better represent similarity between samples, is then applied to SH for effective binary code learning. Meanwhile, our approach, unlike metric learning, can automatically determine the scale factor during the optimization. Extensive experiments are conducted on publicly available datasets and the comparison results demonstrate the effectiveness of our approach.
Meng Wang 0001, Jian Cheng 0001, Changsheng Xu, Hanqing Lu
IEEE Trans. Multim.3
2013 Script-to-Movie: A Computational Framework for Story Movie Composition
abstract
Traditional movie production has always been a highly professional work that needs team collaboration, advanced devices and techniques, and vast time and money investment. These high threshold requirements not only prevent mass amateur enthusiasts entering this field, but also hinder professionals quickly previewing their conceived story plots. In this paper, we raise a novel application, named script-to-movie (S2M) composition, to automatically produce new movies from existing videos in accordance with user created script. Our motivation is to liberate producers from complex filming and editing operations, thereby people's story idea can be instantly converted into the vivid movie video. To support the novel “What You Dream Is What You See” (WYDIWYS) production mode, we first propose a hierarchical alignment method to automatically construct a video material database with detailed semantic description. Considering diverse story plots in user designed script, the database contains abundant video materials about different characters appearing in various time and places conditions. On this basis, the S2M composition is formulated as a constrained optimization problem, where semantic story plot and syntactic visual content are synthetically considered to identify a group of optimal video segments to narrate the user designed script story. Both quantitative and qualitative experimental results are reported to illustrate the effectiveness of the proposed S2M application.
Chao Liang 0001, Changsheng Xu, Jian Cheng 0001, Weiqing Min, Hanqing Lu
IEEE Trans. Multim.3
2012 Modeling Hidden Topics with Dual Local Consistency for Image Analysis
Jian Cheng 0001, Hanqing Lu
ACCV (1)2
2012 Improving Relevance Feedback for Image Retrieval with Asymmetric Sampling
abstract
Relevance feedback is a quite effective approach to improve performance for image retrieval. Recently, active learning method has attracted much attention due to its capability of alleviating the burden of labeling in relevance feedback. However, most of the traditional studies focus on single sample selection in each feedback which needs heavy computational cost in practice. In this paper, we presents a novel batch mode active learning method for informative sample selection. Inspired by graph propagation, we consider the certainty of labels as asymmetric propagation information on graph, and formulate the correlation between labeled samples and unlabeled samples in an united scheme. Extensive experiments on publicly available data sets show that the proposed method is promising.
Biao Niu, Jian Cheng 0001, Hanqing Lu
ICME2
2012 Object detection via foreground contour feature selection and part-based shape model
Huigang Zhang, Junxiu Wang, Xiao Bai 0001, Jun Zhou 0001, Jian Cheng 0001, Huijie Zhao
ICPR5
2012 Dimensionality reduction by Mixed Kernel Canonical Correlation Analysis
Xiaofeng Zhu 0001, Zi Huang, Heng Tao Shen, Jian Cheng 0001, Changsheng Xu
Pattern Recognit.4
2012 Real-Time Probabilistic Covariance Tracking With Efficient Model Update
abstract
The recently proposed covariance region descriptor has been proven robust and versatile for a modest computational cost. The covariance matrix enables efficient fusion of different types of features, where the spatial and statistical properties, as well as their correlation, are characterized. The similarity between two covariance descriptors is measured on Riemannian manifolds. Based on the same metric but with a probabilistic framework, we propose a novel tracking approach on Riemannian manifolds with a novel incremental covariance tensor learning (ICTL). To address the appearance variations, ICTL incrementally learns a low-dimensional covariance tensor representation and efficiently adapts online to appearance changes of the target with only O(1) computational complexity, resulting in a real-time performance. The covariance-based representation and the ICTL are then combined with the particle filter framework to allow better handling of background clutter, as well as the temporary occlusions. We test the proposed probabilistic ICTL tracker on numerous benchmark sequences involving different types of challenges including occlusions and variations in illumination, scale, and pose. The proposed approach demonstrates excellent real-time performance, both qualitatively and quantitatively, in comparison with several previously proposed trackers.
Yi Wu 0001, Jian Cheng 0001, Jinqiao Wang, Hanqing Lu, Haibin Ling, Erik Blasch, Li Bai 0002
IEEE Trans. Image Process.2
2011 TVParser: An automatic TV video parsing method
abstract
In this paper, we propose an automatic approach to simultaneously name faces and discover scenes in TV shows. We follow the multi-modal idea of utilizing script to assist video content understanding, but without using timestamp (provided by script-subtitles alignment) as the connection. Instead, the temporal relation between faces in the video and names in the script is investigated in our approach, and an global optimal video-script alignment is inferred according to the character correspondence. The contribution of this paper is two-fold: (1) we propose a generative model, named TVParser, to depict the temporal character correspondence between video and script, from which face-name relationship can be automatically learned as a model parameter, and meanwhile, video scene structure can be effectively inferred as a hidden state sequence; (2) we find fast algorithms to accelerate both model parameter learning and state inference, resulting in an efficient and global optimal alignment. We conduct extensive comparative experiments on popular TV series and report comparable and even superior performance over existing methods.
Chao Liang 0001, Changsheng Xu, Jian Cheng 0001, Hanqing Lu
CVPR3
2011 Representative sampling with certainty propagation for image retrieval
abstract
Selective sampling has been widely used in relevance feedback of image retrieval to alleviate the burden of labeling by selecting the most informative instances for user to label. Traditional sample selection scheme often selects a batch of instances each time and label them simultaneously, which ignores the correlation among instances and results in redundant labeling. In this paper, we propose an improved representative sampling method with certainty propagation to improve the performance of sampling. In our method, two kinds of correlations among instances are explored to reduce the redundancy in sampling. One is the correlation between labeled instances and unlabeled instances. The other is the correlation among unlabeled instances. Extensive experiments show that the proposed method achieve encouraging results.
Jian Cheng 0001, Biao Niu, Yikai Fang, Hanqing Lu
ICIP1
2011 Learning to detect salient region of image under weak supervision
abstract
Salient region of an image usually contains the crucial information for image analysis and understanding. Most conventional approaches learn the saliency by utilizing the low-level features, which ignore the participation of human. In this paper, we propose an effective and robust approach to detect the salient region of an image by combining the bottom-up and top-down cues. The proposed method not only consider the low-level attention features, but also take human into the loop for better understanding of human attention. Furthermore, we build an asymmetrical graph model to integrate these bottom-up and top-down cues into an energy function of saliency. A compact but exact saliency region can be obtained by minimizing posterior energy function. The compact constraint and global minimization manner of the asymmetrical graph cuts guarantee the good performance of saliency extraction. Extensive experiments demonstrate the proposed method is promising.
Jian Cheng 0001, Hanqing Lu
ICME1
2011 Robust movie character identification and the sensitivity analysis
abstract
Automatic face identification of characters in movies has drawn significant research interests and led to various applications. It is a challenging problem due to the huge variation in the appearance of each character. Although existing methods demonstrate promising results in clean environment, the performances are limited in complex movie scenes due to the noises generated during the face tracking and face clustering process. In this paper we present a robust character identification approach by incorporating a noise insensitive relationship representation and a graph matching algorithm. Beyond existing character identification approaches, we further perform explicit sensitivity analysis on character identification by introducing two types of simulated noises. Experiments validate the advantage of the proposed method.
Chao Liang 0001, Changsheng Xu, Jian Cheng 0001
ICME4
2011 Correlated PLSA for Image Clustering
Jian Cheng 0001, Zechao Li, Hanqing Lu
MMM (1)2
2011 Special edition on semi-supervised learning for visual content analysis and understanding
Jian Cheng 0001, Jinjun Wang, Shuqiang Jiang, Zhi-Hua Zhou, Edwin R. Hancock
Pattern Recognit.1
2010 A co-Gaussian Process based framework for remote sensing image change detection
abstract
Inspired by the idea of co-training algorithm, in this paper we propose a novel semi-supervised learning algorithm, co-Gaussian Process (co-GP), under a Bayesian framework. Image data are characterized in two distinct views, i.e. two disjoint feature sets. A latent function with a GP prior is employed for each view. In learning process of co-GP, knowledge acquired in each view is transferred by probabilistic labels to the other in turns to enhance learning effect. In this manner, proper parameters are estimated in a bootstrap mode and a satisfying performance can be maintained with only small amount of labeled data. The experiments carried out on multitemporal images validate the proposed algorithm.
Zhenglong Li 0001, Jian Cheng 0001, Zhixin Zhou, Hanqing Lu
ICASSP3
2010 A improved silhouette tracking approach integrating particle filter with graph cuts
abstract
In this paper, we propose a novel approach that combines particle filter tracking and 3D graph cut based segmentation to achieve silhouette tracking against drastic scale change and occlusion. The segmentation module offers particle filter tracking procedure the target shape information to compensate spatial information loss in the histogram based particle filter tracking process. Meanwhile, particle filter predicts a location for the shape prior in the segmentation module to overcome the global nature of graph cut algorithm, that is outlying regions similar with object are prone to be captured. The above two parts are linked with a guide mask that is initialized at the beginning of tracking and updated in a mask update mechanism. We demonstrate the effectiveness of our approach with experiments in several challenging image sequences.
Jing Liu 0001, Jinqiao Wang, Jian Cheng 0001, Hanqing Lu
ICASSP4
2010 Personalized Sports Video Customization for Mobile Devices
Chao Liang 0001, Jian Cheng 0001, Changsheng Xu, Jinqiao Wang, Hanqing Lu, Jian Ma 0001
MMM3
2010 Building topographic subspace model with transfer learning for sparse representation
Yang Liu 0021, Jian Cheng 0001, Changsheng Xu, Hanqing Lu
Neurocomputing2
2009 A robust boosting tracker with minimum error bound in a co-training framework
abstract
The varying object appearance and unlabeled data from new frames are always the challenging problem in object tracking. Recently machine learning methods are widely applied to tracking, and some online and semi-supervised algorithms are developed to handle these difficulties. In this paper, we consider tracking as a classification problem and present a novel tracking method based on boosting in a co-training framework. The proposed tracker can be online updated and boosted with multi-view weak hypothesis. The most important contribution of this paper is that we find a boosting error upper bound in a co-training framework to guide the novel tracker construction. In theory, the proposed tracking method is proved to minimize this error bound. In experiments, the accuracy rate of foreground/ background classification and the tracking results are both served as evaluation metrics. Experimental results show good performance of proposed novel tracker on challenging sequences.
Jian Cheng 0001, Hanqing Lu
ICCV2
2009 Real-time visual tracking via Incremental Covariance Tensor Learning
abstract
Visual tracking is a challenging problem, as an object may change its appearance due to pose variations, illumination changes, and occlusions. Many algorithms have been proposed to update the target model using the large volume of available information during tracking, but at the cost of high computational complexity. To address this problem, we present a tracking approach that incrementally learns a low-dimensional covariance tensor representation, efficiently adapting online to appearance changes for each mode of the target with only ̃(1) computational complexity. Moreover, a weighting scheme is adopted to ensure less modeling power is expended fitting older observations. Both of these features contribute measurably to improving overall tracking performance. Tracking is then led by the Bayesian inference framework in which a particle filter is used to propagate sample distributions over time. With the help of integral images, our tracker achieves real-time performance. Extensive experiments demonstrate the effectiveness of the proposed tracking algorithm for the targets undergoing appearance variations.
Yi Wu 0001, Jian Cheng 0001, Jinqiao Wang, Hanqing Lu
ICCV2
2009 Human activity recognition based on the blob features
abstract
In this paper, we present a novel approach for human activities recognition in the video. We analyze human activities in the sequential frames because human activities can be considered as a temporal object which contains a series of frames. Firstly, we establish a statistical background model and extract foreground object through background subtraction in the video stream. Then, we use foreground blobs of the current frame and a series of frames before the current frame to form a new feature image in certain rules. Finally, we combine the non-zero pixels in the feature image into blobs using the connected component method. Then each blob corresponds to an activity which is characterized by the blob appearance. By recognizing blob features we can recognize activities. We use Gaussian Mixture to model features for each type of human activities and employ Mahalanobis distance to measure the similarity.
Jie Yang 0002, Jian Cheng 0001, Hanqing Lu
ICME2
2009 Human-centered picture slideshow personalization for mobile devices
abstract
This paper presents a human-centered picture slideshow system for mobile users. In contrast to conventional ROIs (region-of-interest) detection based systems, we provide mobile users the freedom of personalizing ROIs in a convenient and effective way. Here, we import a simple human interaction, i.e., only a single click, to give a hint for users' ROIs. First, local saliency map (LSM) is generated, which considers not only multi-scale contrast, but also the self-correlation measure and central effect. Then a local fuzzy growing method is adopted to extract ROIs automatically based on LSM. Extensive experiments and user studies show the encouraging performance of the proposed system.
Cunxun Zang, Jian Cheng 0001, Hanqing Lu, Jian Ma 0001
ICME3
2009 Multi-view multi-label active learning for image classification
abstract
Image classification is an important topic in multimedia analysis, among which multi-label image classification is a very challenging task with respect to the large demand for human annotation of multi-label samples. In this paper, we propose a multi-view multi-label active learning strategy, which integrates the mechanism of active learning and multi-view learning. On one hand we explore the sample and label uncertainties within each view; on the other hand we capture the uncertainty over different views based on multi-view fusion. Then the overall uncertainty along the sample, label and view dimensions are obtained to detect the most informative sample-label pairs. Experimental results demonstrate the effectiveness of the proposed scheme.
Xiaoyu Zhang 0002, Jian Cheng 0001, Changsheng Xu, Hanqing Lu, Songde Ma
ICME2
2009 Naming faces in films using hypergraph matching
abstract
In this paper, we aim to address the problem of naming faces in feature-length films using video and film script. Different from the state-of-the-art methods on naming faces in the videos, most of which used a local matching between a visible face and one of the names extracted from the local video transcript, we use a global matching between names and faces as it is not easy to obtain enough local name cues in the films. In the video, we cluster the faces into groups corresponding to characters and build a face network according to face co-occurrence relationship. Similarly in the film script, a name network is also built according to name cooccurrence relationship. The vertices of the two networks are finally matched by a hypergraph matching method. Experiments are conducted on five feature-length films and give encouraging results.
Yifan Zhang 0001, Changsheng Xu, Jian Cheng 0001, Hanqing Lu
ICME3
2009 Semi-supervised Change Detection via Gaussian Processes
abstract
This paper introduces a semi-supervised change detection method that exploits both labeled and unlabeled samples via Gaussian Process (GP). The proposed method is based on recent development in Gaussian Process classifier named NCNM [3]. NCNM is a probabilistic approach to learning a GP classifier in the presence of unlabeled data. It involves a novel transductive learning under a probabilistic framework. Experimental results obtained on two sets of multitemporal remote sensing images confirm the effectiveness of the proposed approach. It also proves that NCNM can compete seriously with the state-of-the-art support vector machines (SVM) classifier for remote sensing image change detection.
Chunlei Huo, Zhixin Zhou, Hanqing Lu, Jian Cheng 0001
IGARSS (2)5
2009 A Variational Bayesian Approach to Remote Sensing Image Change Detection
abstract
In this paper, we present a variational Bayesian (VB) approach to multitemporal remote sensing image change detection. The content of the so called `difference image' is modeled by finite Gaussians Mixture Model (GMM), then with the factor analysis techniques, underlying structure of image content is inferred automatically. Compared with the Expectation-Maximization (EM) algorithm, the proposed method can adaptively determine the number of components in the mixture model without usual sub- or over-segmentation problem. Moreover, to overcome the local optimization problem, a component split strategy is employed in inference process. Experimental results confirm the effectiveness of the proposed method.
Zhenglong Li 0001, Jian Cheng 0001, Zhixin Zhou, Hanqing Lu
IGARSS (3)3
2009 A Variational Co-training Framework for Remote Sensing Image Segmentation
abstract
Inspired by the idea of co-training algorithm, in this paper we propose a novel remote sensing image segmentation approach using co-training strategy under variational Bayesian (VB) framework. Image data are characterized in two distinct views, i.e. two disjoint feature sets. A Gaussian mixture model (GMM) is employed for each view. On one hand, underlying structure of image content is inferred automatically with the factor analysis techniques. On the other hand, parameters are estimated in a bootstrap mode with the co-training strategy. In this manner, a satisfying performance can be achieved. Experimental analyses carried out on several different sets of high resolution optical images validate the proposed algorithm.
Zhenglong Li 0001, Jian Cheng 0001, Zhixin Zhou, Hanqing Lu
IGARSS (4)3
2009 Effective Annotation and Search for Video Blogs with Integration of Context and Content Analysis
abstract
In recent years, weblogs (or blogs) have received great popularity worldwide, among which video blogs (or vlogs) are playing an increasingly important role. However, research on vlog analysis is still in the early stage, and how to manage vlogs effectively so that they can be more easily accessible is a challenging problem. In this paper, we propose a novel vlog management model which is comprised of automatic vlog annotation and user-oriented vlog search. For vlog annotation, we extract informative keywords from both the target vlog itself and relevant external resources; besides semantic annotation, we perform sentiment analysis on comments to obtain the overall evaluation. For vlog search, we present saliency-based matching to simulate human perception of similarity, and organize the results by personalized ranking and category-based clustering. An evaluation criterion is also proposed for vlog annotation, which assigns a score to an annotation according to its accuracy and completeness in representing the vlog's semantics. Experimental results demonstrate the effectiveness of the proposed management model for vlogs.
Xiaoyu Zhang 0002, Changsheng Xu, Jian Cheng 0001, Hanqing Lu, Songde Ma
IEEE Trans. Multim.3
2008 Synchronization analysis for synchronized diving videos
abstract
The judgements in some sports competitions are subjective tasks, especially in competitions with high requirements on skills, which could lead to unfair results. Synchronized diving is such a skillful competition. Using computer to judge competitions automatically or assist referees to judge can effectively decrease the unfair results. This paper presents a framework for analyzing synchronization for synchronized diving videos. The framework is composed of three steps: (1) silhouette extraction, that can effectively get the diverpsilas silhouette; (2) synchronization feature extraction and representation, that can explicitly achieve synchronization features from synchronized diving videos; (3) synchronization evaluation, that can evaluate synchronization with evaluation function which is constructed by statistic learning method with diving video segments come from different games judged by different referees. The experiment shows that the method can accurately and effectively evaluate the synchronization for synchronized diving videos.
Haoyang Ding, Jian Cheng 0001, Hanqing Lu, Zhixin Zhou
ICME2
2008 A novel contextual descriptors for category recognition
abstract
In this paper, we propose a novel contextual descriptor which combines the contextual information and local appearance. Based on Gibbs distribution, a local descriptor is designed. By assembling the contextual information and local descriptors, a new partial contextual descriptor (PCD) is finally presented. Combining Pyramid Match Kernel (PMK) and SVM, we test our new descriptor and obtain higher average precision of classification than using local appearance descriptor.
Ming Tang 0001, Jian Cheng 0001, Jinqiao Wang, Hanqing Lu, Songde Ma
ICME3
2008 Human-centered image navigation on mobile devices
abstract
How to navigate large images on mobile device is an open problem due to small display screen. Currently region of interest (ROI) image compressing is a popular approach. However, since the semantics of image is diversified by different users, it is hard to extract general ROI. In this paper, we propose a human-centered image navigation method for mobile users with simple human interaction. Different from previous work, we first extract Local Saliency Map (LSM) of an image according to personalized requirement, which can reduce the semantic ambiguity of the image for different users. Based on LSM, we can detect the ROIs of the image, which fully satisfies the different interests of the different users. Experiments and user studies show an encouraging performance of the proposed method.
Cunxun Zang, Qingshan Liu 0001, Jian Cheng 0001, Hanqing Lu
ICME3
2008 Automatic semantic annotation for video blogs
abstract
In recent years, Weblogs (or blogs) have received great popularity worldwide, among which video blogs (or vlogs) are playing an increasingly important role. As vlogs gain in population, how to make them more easily accessible has become a hot research topic. In this paper, we propose a novel automatic annotation model for vlogs. We extract informative keywords from both the target vlog itself and external resources which are semantically and visually relevant to it. We also present a new evaluation criterion, which assigns a score to an annotation according to its accuracy and completeness in representing the vlogpsilas semantics. Experimental results demonstrate the effectiveness of both the annotation model and evaluation criterion.
Xiaoyu Zhang 0002, Changsheng Xu, Jian Cheng 0001, Hanqing Lu, Songde Ma
ICME3
2008 Change detection based on adaptive Markov Random Fields
abstract
Usually changes in remote sensing images go along with the appearance or disappearance of some edges. In addition, pixels located along the edges are likely to weakly influenced by its neighborhood pixels, while pixels located far from the edges commonly have a tightly correlation among them. In this paper, we propose a novel change detection technique based on adaptive Markov Random Fields (MRFs) for high resolution satellite images with combined color and texture features. The technique is composed of two main steps: (1) the input images are marked with different region indexes by the combined color and edge features; (2) change maps are obtained under the MRF framework with alterable order of neighborhood and variable smooth weight coefficient controlled by the index map. The main contribution of this paper is that the spatial-contextual information included in the remote sensing imagery is correctly and adaptively exploited under an adaptive MRF framework. Experiments results obtained on a set of remote sensing imagery confirm the effectiveness of the proposed approach.
Chunlei Huo, Jian Cheng 0001, Zhixin Zhou, Hanqing Lu
ICPR3
2008 Hand posture recognition with co-training
abstract
As an emerging human-computer interaction approach vision based hand interaction is more natural and efficient. However in order to achieve high accuracy, most of the existing hand posture recognition methods need a large number of labeled samples which is expensive or unavailable in practice. In this paper, a co-training based method is proposed to recognize different hand postures with a small quantity of labeled data. Hand postures examples are represented with different features and disparate classifiers are trained simultaneously with labeled data. Then the semi-supervised learning treats each new posture as unlabeled data and updates the classifiers in a co-training framework. Experiments show that the proposed method outperforms the traditional methods with much less labeled examples.
Yikai Fang, Jian Cheng 0001, Jinqiao Wang, Kongqiao Wang, Jing Liu 0001, Hanqing Lu
ICPR2
2008 Saliency Cuts: An automatic approach to object segmentation
abstract
Interactive graph cuts are widely used in object segmentation but with some disadvantages: 1) Manual interactions may cause inaccurate or even incorrect segmentation results and involve more interactions especially for novices. 2) In some situations, the manual interactions are infeasible. To overcome these disadvantages, we propose a novel approach, namely Saliency cuts, to segment object from background automatically. By exploring the effects of labels to graph cuts, the so called ldquoprofessional labelsrdquo is introduced to evaluate labels. With the aid of saliency detection, a multiresolution framework is designed to provide ldquoprofessional labelsrdquo automatically and implement object segmentation using graph cuts. The experiments demonstrate the promising performance of Saliency cuts.
Jian Cheng 0001, Zhenglong Li 0001, Hanqing Lu
ICPR2
2008 A variational inference based approach for image segmentation
abstract
In this paper, we present a variational Bayes (VB) approach for image segmentation. First, image is modeled by a mixture model, and then with the techniques of factor analyzer, the underlying structure of image content is inferred automatically. Different from the traditional EM algorithm that seriously suffers from component number selection, the proposed method can accurately infer the underlying image structure including suitable component number without usual sub- or over-segmentation problem. To overcome the problem of local optimization, a component split strategy is adopted in inference optimization process. Extensive experiments on various images validate the proposed method.
Zhenglong Li 0001, Qingshan Liu 0001, Jian Cheng 0001, Hanqing Lu
ICPR3
2008 Multi-cue collaborative kernel tracking with cross ratio invariant constraint
abstract
In this paper, a novel multi-cue collaborative kernel tracking algorithm is proposed. A new constraint based on the property of cross ratio invariant enables tracking of objects insensitive to complex motions, including scale changes, rotation and especially views changes, without labeling and training. Meanwhile, invariant moments are introduced into the kernel based tracking method as the shape representation. The integration of shape and color information makes tracking more robust, and avoids the kernels drifting when color information is not sufficient. Experiments show our method is robust to arbitrary motions of articulated objects and other rigid objects in complex environment.
Jian Cheng 0001, Hanqing Lu
ICPR2
2008 A Multilevel Contextual Approach to Change Detection for very high Resolution Images
abstract
A multilevel contextual approach is proposed in this paper for change detection of VHR images. By representing the change features in a hierarchical contextual manner, the changes are detected level-by-level. By taking advantages of SVMs, the ambiguity of changes is mitigated and the optimal changes are detected peculiar to the specific user. Compared to the traditional methods, the proposed approach is more accurate, more robust and faster. Experiments demonstrate the effectiveness and advantages of the proposed approach.
Chunlei Huo, Zhixin Zhou, Hanqing Lu, Jian Cheng 0001, Qingshan Liu 0001
IGARSS (4)5
2008 Urban Change Detection based on Local Features and Multiscale Fusion
abstract
A multiscale approach is presented in this paper for urban change detection of VHR images. The proposed approach detects the changes at different scales by local-region-based approach, which consists of local region extraction, local region description and local region comparison. To combine the changes at different scales and improve the accuracy, multiscale fusion strategy is applied to local-region-based change detection. Experimental results obtained on Quickbird images confirm the effectiveness of the proposed approach.
Chunlei Huo, Zhixin Zhou, Qingshan Liu 0001, Jian Cheng 0001, Hanqing Lu
IGARSS (3)4
2008 Selective Sampling Based on Dynamic Certainty Propagation for Image Retrieval
Xiaoyu Zhang 0002, Jian Cheng 0001, Hanqing Lu, Songde Ma
MMM2
2007 Image Segmentation Using Co-EM Strategy
Zhenglong Li 0001, Jian Cheng 0001, Qingshan Liu 0001, Hanqing Lu
ACCV (2)2
2007 Robust lip Localization on Multi-View Faces in Video
abstract
In this paper, a fast and robust multi-view lip localization algorithm in video is proposed. We consider lip localization as a binary classification problem, where a classifier is learned to distinguish between the lip and the region surrounding it. The classifier we use here is a histogram-based one which exploits the anthropometrical properties of the human face with the help of face scale normalization. Due to the perceptual uniformity and robustness for lip/skin color variations across different people, we adopt CIELUV color model to represent the color of lip. After classification, we propose a novel projection-cut algorithm by spatial deviation analysis (SDA) to locate the lip, which is effective to deal with the background clutters. Experimental results on teleplay videos demonstrate that the proposed approach is efficient and robust for lip localization.
Yi Wu 0001, Wei Hu 0002, Tao Wang 0003, Yimin Zhang 0002, Jian Cheng 0001, Hanqing Lu
ICIP (4)6
2007 Weighted Co-SVM for Image Retrieval with MVB Strategy
abstract
In relevance feedback, active learning is often used to alleviate the burden of labeling by selecting only the most informative data. Traditional data selection strategies often choose the data closest to the current classification boundary to label, which are in fact not informative enough. In this paper, we propose the moving virtual boundary (MVB) strategy, which is proved to be a more effective way for data selection. The co-SVM algorithm is another powerful method used in relevance feedback. Unfortunately, its basic assumption that each view of the data be sufficient is often untenable in image retrieval. We present our weighted co-SVM as an extension of co-SVM by attaching weight to each view, and thus relax the view sufficiency assumption. The experimental results show that the weighted co-SVM algorithm outperforms co-SVM obviously, especially with the help of MVB data selection strategy.
Xiaoyu Zhang 0002, Jian Cheng 0001, Hanqing Lu, Songde Ma
ICIP (4)2
2007 A Real-Time Hand Gesture Recognition Method
abstract
Compared with the traditional interaction approaches, such as keyboard, mouse, pen, etc, vision based hand interaction is more natural and efficient. In this paper, we proposed a robust real-time hand gesture recognition method. In our method, firstly, a specific gesture is required to trigger the hand detection followed by tracking; then hand is segmented using motion and color cues; finally, in order to break the limitation of aspect ratio encountered in most of learning based hand gesture methods, the scale-space feature detection is integrated into gesture recognition. Applying the proposed method to navigation of image browsing, experimental results show that our method achieves satisfactory performance.
Yikai Fang, Kongqiao Wang, Jian Cheng 0001, Hanqing Lu
ICME3
2007 Active learning for image retrieval with Co-SVM
Jian Cheng 0001, Kongqiao Wang
Pattern Recognit.1
2006 Flexible background mixture models for foreground segmentation
Jian Cheng 0001, Jie Yang 0002, Yue Zhou 0005, Yingying Cui
Image Vis. Comput.1
2006 Ensemble learning for independent component analysis
Jian Cheng 0001, Qingshan Liu 0001, Hanqing Lu, Yen-Wei Chen 0001
Pattern Recognit.1
2005 A Cascaded Ensemble Learning for Independent Component Analysis
Jian Cheng 0001, Kongqiao Wang, Yen-Wei Chen 0001
ISNN (1)1
2005 Supervised kernel locality preserving projections for face recognition
Jian Cheng 0001, Qingshan Liu 0001, Hanqing Lu, Yen-Wei Chen 0001
Neurocomputing1
2004 Real-Time Infrared Object Tracking Based on Mean Shift
Jian Cheng 0001, Jie Yang 0002
CIARP1
2004 A supervised nonlinear local embedding for face recognition
abstract
Many recent works demonstrated that subspace analysis is a good method for face recognition. How to find the subspace is a key issue. In this paper, a supervised nonlinear local embedding (SNLE) method is proposed to construct a subspace for face recognition, in which we combine the idea of nonlinear kernel mapping and preserving local geometric relations of the samples belonging to same class. SNLE can not only gain a perfect approximation of the nonlinear face manifold, but also enhance within-class local information. Moreover, it is also equivalent to solving a generalized eigenvalue problem in mathematics. Our experiments are performed on two benchmarks, and experimental results show that the proposed method has an encouraging performance.
Jian Cheng 0001, Qingshan Liu 0001, Hanqing Lu, Yen-Wei Chen 0001
ICIP1
2004 Random Independent Subspace for Face Recognition
Jian Cheng 0001, Qingshan Liu 0001, Hanqing Lu, Yen-Wei Chen 0001
KES1
2004 Improving ICA Performance for Modeling Image Appearance with the Kernel Trick
Qingshan Liu 0001, Jian Cheng 0001, Hanqing Lu, Songde Ma
KES2
2003 Face Recognition Using Overcomplete Independent Component Analysis
Jian Cheng 0001, Hanqing Lu, Yen-Wei Chen 0001, Xiang-Yan Zeng
KES1
2003 Image Retrieval Based on Independent Components of Color Histograms
Xiang-Yan Zeng, Yen-Wei Chen 0001, Zensho Nakao, Jian Cheng 0001, Hanqing Lu
KES4