VLDB 2026 Research / reviewers in the wild / expert
Bing Deng
dblp:70/3873
· DBLP profile ↗
32ranked-venue papers
3as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 1 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic curriculum knowledge distillation: optimizing knowledge transfer through temporal adaptation
Huaping Zhou, Kelei Sun, Bing Deng |
Multim. Syst. | 6 |
| 2025 | Grounding 3D Object Affordance with Language Instructions, Visual Observations and InteractionsabstractGrounding 3D object affordance is a task that locates objects in 3D space where they can be manipulated, which links perception and action for embodied intelligence. For example, for an intelligent robot, it is necessary to accurately ground the affordance of an object and grasp it according to human instructions. In this paper, we introduce a novel task that grounds 3D object affordance based on language instructions, visual observations and interactions, which is inspired by cognitive science. We collect an Affordance Grounding dataset with Points, Images and Language instructions (AGPIL) to support the proposed task. In the 3D physical world, due to observation orientation, object rotation, or spatial occlusion, we can only get a partial observation of the object. So this dataset includes affordance estimations of objects from full-view, partial-view, and rotation-view perspectives. To accomplish this task, we propose LMAffordance3D, the first multi-modal, language-guided 3D affordance grounding network, which applies a vision-language model to fuse 2D and 3D spatial features with semantic features. Comprehensive experiments on AGPIL demonstrate the effectiveness and superiority of our method on this task, even in unseen experimental settings. Our project is available at https://sites.google.com/view/lmaffordance3d. Quyu Kong, Kechun Xu, Xunlong Xia, Bing Deng, Jieping Ye, Rong Xiong, Yue Wang 0020 |
CVPR | 5 |
| 2025 | PerLDiff: Controllable Street View Synthesis Using Perspective-Layout Diffusion Model
Hualian Sheng, Sijia Cai, Bing Deng, Qiao Liang 0002, Wen Li 0001, Jieping Ye, Shuhang Gu |
ICCV | 4 |
| 2025 | TAU-106K: A New Dataset for Comprehensive Understanding of Traffic AccidentabstractMultimodal Large Language Models (MLLMs) have demonstrated impressive performance in general visual understanding tasks. However, their potential for high-level, fine-grained comprehension, such as anomaly understanding, remains unexplored. Focusing on traffic accidents, a critical and practical scenario within anomaly understanding, we investigate the advanced capabilities of MLLMs and propose TABot, a multimodal MLLM specialized for accident-related tasks. To facilitate this, we first construct TAU-106K, a large-scale multimodal dataset containing 106K traffic accident videos and images collected from academic benchmarks and public platforms. The dataset is meticulously annotated through a video-to-image annotation pipeline to ensure comprehensive and high-quality labels. Building upon TAU-106K, we train TABot using a two-step approach designed to integrate multi-granularity tasks, including accident recognition, spatial-temporal grounding, and an auxiliary description task to enhance the model's understanding of accident elements. Extensive experiments demonstrate TABot's superior performance in traffic accident understanding, highlighting not only its capabilities in high-level anomaly comprehension but also the robustness of the TAU-106K benchmark. Our code and data will be available at https://github.com/cool-xuan/TABot. Yixuan Zhou 0001, Long Bai 0012, Sijia Cai, Bing Deng, Xing Xu 0001, Heng Tao Shen |
ICLR | 4 |
| 2025 | EchoShot: Multi-Shot Portrait Video GenerationabstractVideo diffusion models substantially boost the productivity of artistic workflows with high-quality portrait video generative capacity. However, prevailing pipelines are primarily constrained to single-shot creation, while real-world applications urge multiple shots with identity consistency and flexible content controllability. In this work, we propose EchoShot, a native and scalable multi-shot framework for portrait customization built upon a foundation video diffusion model. To start with, we propose shot-aware position embedding mechanisms within the video diffusion transformer architecture to model inter-shot variations and establish intricate correspondence between multi-shot visual content and their textual descriptions. This simple yet effective design enables direct training on multi-shot video data without introducing additional computational overhead. To facilitate model training within multi-shot scenarios, we construct PortraitGala, a large-scale and high-fidelity human-centric video dataset featuring cross-shot identity consistency and fine-grained captions such as facial attributes, outfits, and dynamic motions. To further enhance applicability, we extend EchoShot to perform reference image-based personalized multi-shot generation and long video synthesis with infinite shot counts. Extensive evaluations demonstrate that EchoShot achieves superior identity consistency as well as attribute-level controllability in multi-shot portrait video generation. Notably, the proposed framework demonstrates potential as a foundational paradigm for general multi-shot video modeling. Project page: https://johnneywang.github.io/EchoShot-webpage. Jiahao Wang 0004, Hualian Sheng, Sijia Cai, Weizhan Zhang, Caixia Yan, Yachuang Feng, Bing Deng, Jieping Ye |
NeurIPS | 7 |
| 2025 | CT3D++: Improving 3D Object Detection with Keypoint-Induced Channel-wise Transformer
Hualian Sheng, Sijia Cai, Na Zhao 0004, Bing Deng, Qiao Liang 0002, Minjian Zhao, Jieping Ye |
Int. J. Comput. Vis. | 4 |
| 2025 | Efficient Alignment of Unconditioned Action Prior for Language-Conditioned Pick and Place in ClutterabstractWe study the task of language-conditioned pick and place in clutter, where a robot should grasp a target object in open clutter and move it to a specified place. Some approaches learn end-to-end policies with features from vision foundation models, requiring large datasets. Others combine foundation models in a zero-shot setting, suffering from cascading errors. In addition, they primarily leverage vision and language foundation models, focusing less on action priors. In this paper, we aim to develop an effective policy by integrating foundation priors from vision, language, and action. We propose A2, an action prior alignment method that aligns unconditioned action priors with 3D vision-language priors by learning one attention layer. The alignment formulation enables our policy to train with less data and preserve zero-shot generalization capabilities. We show that a shared policy for both pick and place actions enhances the performance for each task, and introduce a policy adaptation scheme to accommodate the multi-modal nature of actions. Extensive experiments in simulation and the real-world show that our policy achieves higher task success rates with fewer steps for both pick and place tasks in clutter, effectively generalizing to unseen objects and language instructions. Videos and codes are available at https://xukechun.github.io/papers/A2. Kechun Xu, Xunlong Xia, Yunxuan Mao, Bing Deng, Jieping Ye, Rong Xiong, Yue Wang 0020 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | Sudoku: Scalable High-Density Cloud Rendering Multi-Client Architecture
Yun Wang 0039, Bing Deng, Xia Jiang, Xuyan Hu, Dongjie Tang, Randy Xu, Yijin Sun, Zhengwei Qi |
IEEE Trans. Serv. Comput. | 2 |
| 2024 | RoScenes: A Large-Scale Multi-view 3D Dataset for Roadside Perception
Xiaosu Zhu, Hualian Sheng, Sijia Cai, Bing Deng, Shaopeng Yang, Qiao Liang 0002, Ken Chen 0005, Lianli Gao, Jingkuan Song, Jieping Ye |
ECCV (41) | 4 |
| 2024 | OTVIC: A Dataset with Online Transmission for Vehicle-to-Infrastructure Cooperative 3D Object DetectionabstractVehicle-to-infrastructure cooperative 3D object detection (VIC3D) is a task that leverages both vehicle and roadside sensors to jointly perceive the surrounding environment. However, considering the high speed of vehicles, the real-time requirements, and the limitations of communication bandwidth, roadside devices transmit the results of perception rather than raw sensor data or feature maps in our real-world scenarios. And affected by various environmental factors, the transmission delay is dynamic. To meet the needs of practical applications, we present OTVIC, which is the first multi-modality and multi-view dataset with online transmission from real scenes for vehicle-to-infrastructure cooperative 3D object detection. The ego-vehicle receives the results of infrastructure perception in real-time, collected from a section of highway in Chengdu, China. Moreover, we propose LfFormer, which is a novel end-to-end multi-modality late fusion framework with transformer for VIC3D task as a baseline based on OTVIC. Experiments prove our fusion framework’s effectiveness and robustness. Our project is available at https://sites.google.com/view/otvic. Yunkai Wang, Quyu Kong, Yufei Wei, Xunlong Xia, Bing Deng, Rong Xiong, Yue Wang 0020 |
IROS | 6 |
| 2024 | Versatile correlation learning for size-robust generalized counting: A new perspective
Hanqing Yang 0002, Sijia Cai, Bing Deng, Mohan Wei, Yu Zhang 0018 |
Knowl. Based Syst. | 3 |
| 2024 | Context-Aware and Semantic-Consistent Spatial Interactions for One-Shot Object Detection Without Fine-TuningabstractOne-shot object detection (OSOD) without fine-tuning has recently garnered considerable attention and research focus. It aims to directly detect novel-class objects in the target image by providing merely one support image patch without undergoing the fine-tuning stage. However, most existing methods adopt image pair matching regardless of the scale inconsistency and spatial semantic mismatch of image pairs, which limits their ability to acquire high-quality target-support related features. This paper addresses these limitations by incorporating cross-scale contexts and semantic-consistent cues that are robust against the challenges of scarce and ambiguous matching. Specifically, we first introduce a simple yet effective Aggregation-Transformer-based Pyramid (ATP) module to explore the long-range cross-scale spatial interactions by employing the customized size-aware aggregation approach and the vanilla transformer encoder, thus the coarse-to-fine local image patterns are optimally utilized. Furthermore, we formulate the 4D contrastive cross-correlation tensor for instance-level features matching and suggest a Geometric Consistent Correlation (GCC) module that utilizes the bidirectional spatial-aware convolutions to extract the long-range semantic correspondences for target-support pairs. Additionally, a Channel Contrastive Learning (CCL) branch is adopted to complement the inter-channel interactions between target-support pairs for the GCC module. Extensive experiments demonstrate that our approach significantly outperforms the previous state-of-the-art methods by 6.5% and 2.1% on PASCAL VOC and COCO datasets for unseen classes, respectively. Hanqing Yang 0002, Sijia Cai, Bing Deng, Jieping Ye, Guosheng Lin, Yu Zhang 0018 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Bindox: An Efficient and Secure Cross-System IPC Mechanism for Multi-Platform ContainersabstractContainerization is widely used for isolation in various applications because it is lightweight, scalable, and portable.In modern distributed systems, seamless inter-process communication (IPC) between multi-platform containers is essential for a range of applications and services, including microservices, cloud computing, and Internet of Things (IoT) devices.However, secure and efficient communication between containers on the same host is challenging, especially when different operating systems are involved.This paper introduces Bindox, a lightweight, efficient, and secure IPC mechanism that enables seamless communication across multiple platforms, including Android and Linux.Bindox uses shared memory for data transfer and implements a stable client-server architecture, ensuring high performance and ease of maintenance.Additionally, Bindox provides a robust security mechanism that guarantees confidentiality, integrity, and availability of the communication channel.Experimental results demonstrate that Bindox outperforms existing networking and IPC methods in terms of memory use, latency, and CPU usage, making it a promising solution for efficient and secure communication between multi-platform containers. Yuxin Xiang, Bing Deng, Randy Xu, Marc Mao, Yun Wang 0039, Zhengwei Qi |
SEKE | 2 |
| 2023 | A novel feature and sample joint transfer learning method with feature selection in semi-supervised scenarios for identifying the sequence of some species with less known genetic data
Jianghui Wen, Haoran Huang, Zhenyu Pu, Bing Deng |
Soft Comput. | 4 |
| 2023 | PDR: Progressive Depth Regularization for Monocular 3D Object DetectionabstractAccurately predicting object depth is a key challenge in monocular 3D detection task. The perspective projection principle used by most state-of-the-art approaches demands a complex balance between the ratio-form depth estimation and 2D-3D geometric regularizations, and thus can lead to sub-optimal solutions. In this paper, we propose a novel synergistic scheme that can achieve better trade-off among these competing objectives. Our main proposal is a progressive depth regularization (PDR) architecture that splits the overall training process into three sequential depth estimation steps to gradually remove the unwanted deviations induced by the over-regularization. Specifically, our model first learns the coarse depth with the conventional perspective projection and combines the coarse-to-fine generation to reduce the search space of 2D projection height prediction. We then deactivate individual supervision on 2D projection height prediction and introduces a new auxiliary 3D physical height prediction to relax the 2D and 3D regularizations, respectively. Consequently, our PDR leads to more precise depth estimation by mitigating the inherent ambiguities in the geometric priors of perspective projection through progressive regularization relaxation. Extensive experiments on both KITTI and Rope3D benchmark show that our PDR delivers strong performance gains as compared to the previous methods. Hualian Sheng, Sijia Cai, Na Zhao 0004, Bing Deng, Minjian Zhao, Gim Hee Lee |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Balanced and Hierarchical Relation Learning for One-shot Object DetectionabstractInstance-level feature matching is significantly important to the success of modern one-shot object detectors. Re-cently, the methods based on the metric-learning paradigm have achieved an impressive process. Most of these works only measure the relations between query and target objects on a single level, resulting in suboptimal performance overall. In this paper, we introduce the balanced and hierarchical learning for our detector. The contributions are two-fold: firstly, a novel Instance-level Hierarchical Relation (IHR) module is proposed to encode the contrastive-level, salient-level, and attention-level relations simultane-ously to enhance the query-relevant similarity representation. Secondly, we notice that the batch training of the IHR module is substantially hindered by the positive-negative sample imbalance in the one-shot scenario. We then in-troduce a simple but effective Ratio-Preserving Loss (RPL) to protect the learning of rare positive samples and sup-press the effects of negative samples. Our loss can adjust the weight for each sample adaptively, ensuring the desired positive-negative ratio consistency and boosting query-related IHR learning. Extensive experiments show that our method outperforms the state-of-the-art method by 1.6% and 1.3% on PASCAL VOC and MS COCO datasets for unseen classes, respectively. The code will be available at https://github.com/hero-y/BHRL. Hanqing Yang 0002, Sijia Cai, Hualian Sheng, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Yu Zhang 0018 |
CVPR | 4 |
| 2022 | Rethinking IoU-based Optimization for Single-stage 3D Object Detection
Hualian Sheng, Sijia Cai, Na Zhao 0004, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Minjian Zhao, Gim Hee Lee |
ECCV (9) | 4 |
| 2022 | Centerness-Aware Network for Temporal Action ProposalabstractTemporal action proposal generation aims at localizing the temporal segments containing human actions in a video. This work proposes a centerness-aware network (CAN), which is a novel one-stage approach intended to generate action proposals as keypoint triplets. A keypoint triplet contains two boundary points (starting and ending) and one center point. Specifically, we evaluate the probabilities of each temporal location in the video whether it is at the boundaries or the center region of ground truth action proposals. CAN optimizes the predicted boundary points interactively in a bidirectional adaptation form by exploiting the dependencies among them. Furthermore, to accurately locate the center points of action proposals with different time spans, temporal feature pyramids are utilized to incorporate multi-scale information explicitly. Using the generated three keypoints, CAN efficiently retrieves temporal proposals by grouping keypoints into triplets if they are geometrically aligned. Experiments show that CAN achieves the state-of-the-art performance on the public THUMOS-14 and ActivityNet-1.3 datasets. Moreover, further experiments demonstrate that by applying action classifiers on proposals generated by CAN, our method achieves the state-of-the-art performance in temporal action localization. Yuan Liu 0017, Jingyuan Chen 0003, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Revisiting Knowledge Distillation: An Inheritance and Exploration FrameworkabstractKnowledge Distillation (KD) is a popular technique to transfer knowledge from a teacher model or ensemble to a student model. Its success is generally attributed to the privileged information on similarities/consistency between the class distributions or intermediate feature representations of the teacher model and the student model. However, directly pushing the student model to mimic the probabilities/features of the teacher model to a large extent limits the student model in learning undiscovered knowledge/features. In this paper, we propose a novel inheritance and exploration knowledge distillation framework (IE-KD), in which a student model is split into two parts - inheritance and exploration. The inheritance part is learned with a similarity loss to transfer the existing learned knowledge from the teacher model to the student model, while the exploration part is encouraged to learn representations different from the inherited ones with a dis-similarity loss. Our IE-KD framework is generic and can be easily combined with existing distillation or mutual learning methods for training deep neural networks. Extensive experiments demonstrate that these two parts can jointly push the student model to learn more diversified and effective representations, and our IE-KD can be a general technique to improve the student network to achieve SOTA performance. Furthermore, by applying our IE-KD to the training of two networks, the performance of both can be improved w.r.t. deep mutual learning. Zhen Huang 0007, Xu Shen 0001, Jun Xing, Tongliang Liu, Xinmei Tian 0001, Houqiang Li, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001 |
CVPR | 7 |
| 2021 | The Blessings of Unlabeled Background in Untrimmed VideosabstractWeakly-supervised Temporal Action Localization (WTAL) aims to detect the action segments with only video-level action labels in training. The key challenge is how to distinguish the action of interest segments from the background, which is unlabelled even on the video-level. While previous works treat the background as "curses", we consider it as "blessings". Specifically, we first use causal analysis to point out that the common localization errors are due to the unobserved confounder that resides ubiquitously in visual recognition. Then, we propose a Temporal Smoothing PCA-based (TS-PCA) deconfounder, which exploits the unlabelled background to model an observed substitute for the unobserved confounder, to remove the confounding effect. Note that the proposed deconfounder is model-agnostic and non-intrusive, and hence can be applied in any WTAL method without model re-designs. Through extensive experiments on four state-of-the-art WTAL methods, we show that the deconfounder can improve all of them on the public datasets: THUMOS-14 and ActivityNet-1.31. Yuan Liu 0002, Jingyuan Chen 0003, Zhenfang Chen, Bing Deng, Jianqiang Huang 0001, Hanwang Zhang |
CVPR | 4 |
| 2021 | DCT-Mask: Discrete Cosine Transform Mask Representation for Instance SegmentationabstractBinary grid mask representation is broadly used in instance segmentation. A representative instantiation is Mask R-CNN which predicts masks on a 28×28 binary grid. Generally, a low-resolution grid is not sufficient to capture the details, while a high-resolution grid dramatically increases the training complexity. In this paper, we propose a new mask representation by applying the discrete cosine transform(DCT) to encode the high-resolution binary grid mask into a compact vector. Our method, termed DCT-Mask, could be easily integrated into most pixel-based instance segmentation methods. Without any bells and whistles, DCT-Mask yields significant gains on different frameworks, backbones, datasets, and training schedules. It does not require any pre-processing or pre-training, and almost no harm to the running speed. Especially, for higher-quality annotations and more complex backbones, our method has a greater improvement. Moreover, we analyze the performance of our method from the perspective of the quality of mask representation. The main reason why DCT-Mask works well is that it obtains a high-quality mask representation with low complexity. Jirui Yang, Chunbo Wei, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Kewei Liang |
CVPR | 4 |
| 2021 | Enhanced Fixed-Interval Smoothing for Markovian Switching Systems
Xi Li 0020, Le Yang 0001, Lyudmila Mihaylova, Bing Deng |
FUSION | 5 |
| 2021 | Improving 3D Object Detection with Channel-wise TransformerabstractThough 3D object detection from point clouds has achieved rapid progress in recent years, the lack of flexible and high-performance proposal refinement remains a great hurdle for existing state-of-the-art two-stage detectors. Previous works on refining 3D proposals have relied on human-designed components such as keypoints sampling, set abstraction and multi-scale feature fusion to produce powerful 3D object representations. Such methods, however, have limited ability to capture rich contextual dependencies among points. In this paper, we leverage the high-quality region proposal network and a Channel-wise Transformer architecture to constitute our two-stage 3D object detection framework (CT3D) with minimal hand-crafted design. The proposed CT3D simultaneously performs proposal-aware embedding and channel-wise context aggregation for the point features within each proposal. Specifically, CT3D uses proposal’s keypoints for spatial contextual modelling and learns attention propagation in the encoding module, mapping the proposal to point embeddings. Next, a new channel-wise decoding module enriches the query-key interaction via channel-wise re-weighting to effectively merge multi-level contexts, which contributes to more accurate object predictions. Extensive experiments demonstrate that our CT3D method has superior performance and excellent scalability. Remarkably, CT3D achieves the AP of 81.77% in the moderate car category on the KITTI test 3D detection benchmark, outperforms state-of-the-art 3D detectors. Hualian Sheng, Sijia Cai, Yuan Liu 0017, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001, Minjian Zhao |
ICCV | 4 |
| 2019 | Quantization NetworksabstractAlthough deep neural networks are highly effective, their high computational and memory costs severely hinder their applications to portable devices. As a consequence, lowbit quantization, which converts a full-precision neural network into a low-bitwidth integer version, has been an active and promising research topic. Existing methods formulate the low-bit quantization of networks as an approximation or optimization problem. Approximation-based methods confront the gradient mismatch problem, while optimizationbased methods are only suitable for quantizing weights and can introduce high computational cost during the training stage. In this paper, we provide a simple and uniform way for weights and activations quantization by formulating it as a differentiable non-linear function. The quantization function is represented as a linear combination of several Sigmoid functions with learnable biases and scales that could be learned in a lossless and end-to-end manner via continuous relaxation of the steepness of Sigmoid functions. Extensive experiments on image classification and object detection tasks show that our quantization networks outperform state-of-the-art methods. We believe that the proposed method will shed new lights on the interpretation of neural network quantization. Jiwei Yang, Xu Shen 0001, Jun Xing, Xinmei Tian 0001, Houqiang Li, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001 |
CVPR | 6 |
| 2019 | A classification model for lncRNA and mRNA based on k-mers and a convolutional neural networkabstractBACKGROUND: Long-chain non-coding RNA (lncRNA) is closely related to many biological activities. Since its sequence structure is similar to that of messenger RNA (mRNA), it is difficult to distinguish between the two based only on sequence biometrics. Therefore, it is particularly important to construct a model that can effectively identify lncRNA and mRNA. RESULTS: First, the difference in the k-mer frequency distribution between lncRNA and mRNA sequences is considered in this paper, and they are transformed into the k-mer frequency matrix. Moreover, k-mers with more species are screened by relative entropy. The classification model of the lncRNA and mRNA sequences is then proposed by inputting the k-mer frequency matrix and training the convolutional neural network. Finally, the optimal k-mer combination of the classification model is determined and compared with other machine learning methods in humans, mice and chickens. The results indicate that the proposed model has the highest classification accuracy. Furthermore, the recognition ability of this model is verified to a single sequence. CONCLUSION: We established a classification model for lncRNA and mRNA based on k-mers and the convolutional neural network. The classification accuracy of the model with 1-mers, 2-mers and 3-mers was the highest, with an accuracy of 0.9872 in humans, 0.8797 in mice and 0.9963 in chickens, which is better than those of the random forest, logistic regression, decision tree and support vector machine. Jianghui Wen, Yeshu Liu, Haoran Huang, Bing Deng |
BMC Bioinform. | 5 |
| 2017 | Stylized Adversarial AutoEncoder for Image GenerationabstractIn this paper, we propose an autoencoder-based generative adversarial network (GAN) for automatic image generation, which is called "stylized adversarial autoencoder". Different from existing generative autoencoders which typically impose a prior distribution over the latent vector, the proposed approach splits the latent variable into two components: style feature and content feature, both encoded from real images. The split of the latent vector enables us adjusting the content and the style of the generated image arbitrarily by choosing different exemplary images. In addition, a multiclass classifier is adopted in the GAN network as the discriminator, which makes the generated images more realistic. We performed experiments on hand-writing digits, scene text and face datasets, in which the stylized adversarial autoencoder achieves superior results for image generation as well as remarkably improves the corresponding supervised recognition task. Yiru Zhao, Bing Deng, Jianqiang Huang 0001, Hongtao Lu 0001, Xian-Sheng Hua 0001 |
ACM Multimedia | 2 |
| 2017 | Spatio-Temporal AutoEncoder for Video Anomaly DetectionabstractAnomalous events detection in real-world video scenes is a challenging problem due to the complexity of "anomaly" as well as the cluttered backgrounds, objects and motions in the scenes. Most existing methods use hand-crafted features in local spatial regions to identify anomalies. In this paper, we propose a novel model called Spatio-Temporal AutoEncoder (ST AutoEncoder or STAE), which utilizes deep neural networks to learn video representation automatically and extracts features from both spatial and temporal dimensions by performing 3-dimensional convolutions. In addition to the reconstruction loss used in existing typical autoencoders, we introduce a weight-decreasing prediction loss for generating future frames, which enhances the motion feature learning in videos. Since most anomaly detection datasets are restricted to appearance anomalies or unnatural motion anomalies, we collected a new challenging dataset comprising a set of real-world traffic surveillance videos. Several experiments are performed on both the public benchmarks and our traffic dataset, which show that our proposed method remarkably outperforms the state-of-the-art approaches. Yiru Zhao, Bing Deng, Chen Shen 0003, Yao Liu 0014, Hongtao Lu 0001, Xian-Sheng Hua 0001 |
ACM Multimedia | 2 |
| 2010 | Comments on "A Convolution and Product Theorem for the Linear Canonical Transform"abstractA recent letter proposed a new convolution structure for the linear canonical transform, and claimed that theirs was clearly easier to implement in the designing of filters than the one suggested earlier in. However, we find that the two kinds of filtering methods are essentially the same through the theoretic deduction. That is to say, the two kinds of filtering methods can obtain the same effect through the same filtering steps. Taking digital signal processing into account, we analyze further the computation complexity according to the steps of multiplicative filter. Bing Deng, Ran Tao 0003, Yue Wang 0001 |
IEEE Signal Process. Lett. | 1 |
| 2006 | Convolution theorems for the linear canonical transform and their applications
Bing Deng, Ran Tao 0003, Yue Wang 0001 |
Sci. China Ser. F Inf. Sci. | 1 |
| 2006 | Research progress of the fractional Fourier transform in signal processing
Ran Tao 0003, Bing Deng, Yue Wang 0001 |
Sci. China Ser. F Inf. Sci. | 2 |
| 2004 | Nonlinear Dynamic Method to Suppress Reverberation Based on RBF Neural Networks
Bing Deng, Ran Tao 0003 |
ISNN (2) | 1 |
| 2004 | Quality Improvement Modeling and Practice in Baosteel Based on KIV-KOV Analysis
Haidong Tang, Zhiping Fan, Bing Deng |
WAIM | 4 |