EDBT 2026 Demo / reviewers in the wild / expert
Huicheng Zheng
dblp:00/3034
· DBLP profile ↗
60ranked-venue papers
10as first author
28since 2021 · last 2026
0000-0002-6729-4176ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 4 first-author · 16 since 2021Artificial intelligence and machine learning · 29 · 7 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Elevating descriptive excellence: Object-centric dense video captioning
Zehua Liu, Huicheng Zheng, Yun Lan |
Neurocomputing | 2 |
| 2026 | Fast Track Anything With Sparse Spatio-Temporal Propagation for Unified Video SegmentationabstractRecent advances in "track-anything" models have significantly improved fine-grained video understanding by simultaneously handling multiple video segmentation and tracking tasks. However, existing models often struggle with robust and efficient temporal propagation. To address these challenges, we propose the Sparse Spatio-Temporal Propagation (SSTP) method, which achieves robust and efficient unified video segmentation by selectively leveraging key spatio-temporal features in videos. Specifically, we design a dynamic 3D spatio-temporal convolution to aggregate global multi-frame spatio-temporal information into memory frames during memory construction. Additionally, we introduce a spatio-temporal aggregation reading strategy to efficiently aggregate the relevant spatio-temporal features from multiple memory frames during memory retrieval. By combining SSTP with an image segmentation foundation model, such as the segment anything model, our method effectively addresses multiple data-scarce video segmentation tasks. Our experimental results demonstrate state-of-the-art performance on five video segmentation tasks across eleven datasets, outperforming both task-specific and unified methods. Notably, SSTP exhibits strong robustness in handling sparse, low-frame-rate videos, making it well-suited for real-world applications. Jisheng Dang, Huicheng Zheng, Zhixuan Chen, Yulan Guo, Tat-Seng Chua |
IEEE Trans. Image Process. | 2 |
| 2026 | Video Decoupling Networks for Accurate, Efficient, Generalizable, and Robust Video Object SegmentationabstractVideo object segmentation (VOS) is a fundamental task in video analysis, aiming to accurately recognize and segment objects of interest within video sequences. Conventional methods, relying on memory networks to store single-frame appearance features, face challenges in computational efficiency and capturing dynamic visual information effectively. To address these limitations, we present a Video Decoupling Network (VDN) with a per-clip memory updating mechanism. Our approach is inspired by the dual-stream hypothesis of the human visual cortex and decomposes multiple previous video frames into fundamental elements: scene, motion, and instance. We propose the Unified Prior-based Spatio-temporal Decoupler (UPSD) algorithm, which parses multiple frames into basic elements in a unified manner. UPSD continuously stores elements over time, enabling adaptive integration of different cues based on task requirements. This decomposition mechanism facilitates comprehensive spatial-temporal information capture and rapid updating, leading to notable enhancements in overall VOS performance. Extensive experiments conducted on multiple VOS benchmarks validate the state-of-the-art accuracy, efficiency, generalizability, and robustness of our approach. Remarkably, VDN demonstrates a significant performance improvement and a substantial speed-up compared to previous state-of-the-art methods on multiple VOS benchmarks. It also exhibits excellent generalizability under domain shift and robustness against various noise types. Jisheng Dang, Huicheng Zheng, Yulan Guo, Jian-Huang Lai, Bin Hu 0001, Tat-Seng Chua |
IEEE Trans. Image Process. | 2 |
| 2025 | External Memory Matters: Generalizable Object-Action Memory for Retrieval-Augmented Long-Term Video UnderstandingabstractLong video understanding with Large Language Models (LLMs) enables the description of objects that are not explicitly present in the training data. However, continuous changes in known objects and the emergence of new ones require up-to-date knowledge of objects and their dynamics for effective understanding of the open world. To alleviate this, we propose an efficient Retrieval-Enhanced Video Understanding method, dubbed REVU, which leverages external knowledge to enhance the performance of open-world learning. First, REVU introduces an extensible external text-object memory with minimal text-visual mapping, involving static and dynamic multimodal information to help LLMs-based models align text and vision features. Second, REVU retrieves object information from external databases and dynamically integrates frame-specific data from videos, enabling effective knowledge aggregation to comprehend the open world. We conducted experiments on multiple benchmark datasets, and our model demonstrates strong adaptability to out-of-domain data without requiring additional fine-tuning or re-training. Experiments on benchmark video understanding datasets reveal that our model achieves state-of-the-art performance and robust generalization. Jisheng Dang, Huicheng Zheng, Jingmei Jiao, Bimei Wang, Bin Hu 0001, Jian-Huang Lai, Tat-Seng Chua |
IJCAI | 2 |
| 2025 | DIGL: Domain-Invariant Global-Local Feature Learning for Gaze Estimation
Wu He, Huicheng Zheng, Jiyuan Lin |
PRCV (2) | 3 |
| 2025 | Adaptive Sparse Memory Networks for Efficient and Robust Video Object SegmentationabstractRecently, memory-based networks have achieved promising performance for video object segmentation (VOS). However, existing methods still suffer from unsatisfactory segmentation accuracy and inferior efficiency. The reasons are mainly twofold: 1) during memory construction, the inflexible memory storage mechanism results in a weak discriminative ability for similar appearances in complex scenarios, leading to video-level temporal redundancy, and 2) during memory reading, matching robustness and memory retrieval accuracy decrease as the number of video frames increases. To address these challenges, we propose an adaptive sparse memory network (ASM) that efficiently and effectively performs VOS by sparsely leveraging previous guidance while attending to key information. Specifically, we design an adaptive sparse memory constructor (ASMC) to adaptively memorize informative past frames according to dynamic temporal changes in video frames. Furthermore, we introduce an attentive local memory reader (ALMR) to quickly retrieve relevant information using a subset of memory, thereby reducing frame-level redundant computation and noise in a simpler and more convenient manner. To prevent key features from being discarded by the subset of memory, we further propose a novel attentive local feature aggregation (ALFA) module, which preserves useful cues by selectively aggregating discriminative spatial dependence from adjacent frames, thereby effectively increasing the receptive field of each memory frame. Extensive experiments demonstrate that our model achieves state-of-the-art performance with real-time speed on six popular VOS benchmarks. Furthermore, our ASM can be applied to existing memory-based methods as generic plugins to achieve significant performance improvements. More importantly, our method exhibits robustness in handling sparse videos with low frame rates. Jisheng Dang, Huicheng Zheng, Xiaohao Xu, Longguang Wang, Qingyong Hu, Yulan Guo |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Confidence-Guided Feature Alignment for Cloth-Changing Person Re-identification
Sirong Huang, Huicheng Zheng |
ICPR (29) | 2 |
| 2024 | An Instance-Level Motion-Aware Graph Model for Multi-target Multi-camera TrackingabstractMulti-target multi-camera tracking (MTMCT), focusing on inferring trajectory across multiple surveillance videos, is holding significant practical utility. While numerous studies aim to learn visual features robust to illumination variation, occlusions, and other issues, the spatio-temporal information within a multi-camera system is not explored sufficiently. Existing methods generally exploit coarse-grained spatio-temporal information by modeling the distribution of time intervals between cameras. However, they tend to neglect the specific motion state of individual objects, which may hinder accurate cross-camera association. In this paper, we introduce a novel motionaware graph (MAG) model designed to extract instance-level spatio-temporal information that aligns with visual features and seamlessly aggregate these two kinds of information within a unified graph framework for MTMCT. Specifically, we propose a motion encoder-decoder module that predicts spatio-temporal consistency scores between objects based on their instance-level motion states. These scores are then integrated with visual similarity scores to generate discriminative feature representations for data association via a graph attention mechanism. Experimental evaluations and ablation studies on the large-scale MTA dataset demonstrate the superiority of our proposed model. Xiaotong Fan, Huicheng Zheng |
IJCNN | 2 |
| 2024 | Histogram Prediction and Equalization for Indoor Monocular Depth Estimation
Bojie Chen, Huicheng Zheng |
PRCV (3) | 2 |
| 2024 | Reversible Data Hiding With Pattern Adaptive PredictionabstractAbstract In the area of reversible data hiding (RDH), one of the most popular techniques is prediction-error expansion (PEE), which hides data in the prediction errors with well-preserved image fidelity. The key to a successful PEE-based RDH implementation usually lies in prediction algorithms with high accuracy. Existing PEE-based RDH works often employ one single prediction algorithm, which is usually globally optimized, but with less consideration of the pixel distribution characteristics within local neighborhoods. In this manuscript, the technique of pattern adaptive prediction is proposed for pixel estimation according to the type of local binary pattern (LBP), which is obtained from the pixel’s eight neighborhood. Theoretically speaking, pattern-based predictors can be designed for each and every LBP patterns to create multiple prediction-error histograms (PEHs). However, the process of performance optimization with multiple PEHs requires extremely heavy computing power. To speed up the optimization process, LBP patterns are classified into various groups based on the degree of histogram concentration. Experiments demonstrate that the prediction accuracy is obviously improved and the image fidelity is well preserved. Junying Yuan, Huicheng Zheng, Jiangqun Ni |
Comput. J. | 2 |
| 2024 | Beyond Appearance: Multi-Frame Spatio-Temporal Context Memory Networks for Efficient and Robust Video Object SegmentationabstractCurrent video object segmentation approaches primarily rely on frame-wise appearance information to perform matching. Despite significant progress, reliable matching becomes challenging due to rapid changes of the object's appearance over time. Moreover, previous matching mechanisms suffer from redundant computation and noise interference as the number of accumulated frames increases. In this paper, we introduce a multi-frame spatio-temporal context memory (STCM) network to exploit discriminative spatio-temporal cues in multiple adjacent frames by utilizing a multi-frame context interaction module (MCI) for memory construction. Based on the proposed MCI module, a sparse group memory reader is developed to enable efficient sparse matching during memory reading. Our proposed method is generic and achieves state-of-the-art performance with real-time speed on benchmark datasets such as DAVIS and YouTube-VOS. In addition, our model exhibits robustness to sparse videos with low frame rates. Jisheng Dang, Huicheng Zheng, Xiaohao Xu, Longguang Wang, Yulan Guo |
IEEE Trans. Image Process. | 2 |
| 2024 | Temporo-Spatial Parallel Sparse Memory Networks for Efficient Video Object SegmentationabstractMemory-based networks have achieved tremendous success in video object segmentation. However, these methods still suffer from unfaithful segmentation and inferior efficiency under complicated video scenarios. The reasons are mainly threefold: 1) Weak perception of fast-moving targets due to individual frame memory patterns without capturing inter-frame motion; 2) Lack of discrimination to visually similar appearances due to the limited receptive field; 3) Redundant computation caused by matching with all memorized frames. To address these issues, we propose a Temporo-Spatial Parallel Sparse Memory network (TSPSM) for efficient video object segmentation. Our TSPSM constructs a temporal memory bank and a spatial memory bank in parallel to memorize complementary discriminative object cues. The temporal bank exploits discriminative temporal motion cues, while the spatial bank mines spatial context cues between adjacent frames with large receptive fields, thereby alleviating the ambiguity caused by similar instances and fast movements. To reduce redundant computation without sacrificing performance during the matching step, we further design a parallel sparse memory reader based on the constructed informative memory banks, which efficiently retrieves relevant temporal and spatial information in a parallel way. Experiments demonstrate that our TSPSM achieves state-of-the-art performance with real-time speed on DAVIS, and YouTube-VOS benchmarks. Furthermore, extensive experiments show that the proposed TSPMC module can be applied to existing methods as a generic plugin to significantly improve performance. Jisheng Dang, Huicheng Zheng, Bimei Wang, Longguang Wang, Yulan Guo |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Unified Spatio-Temporal Dynamic Routing for Efficient Video Object SegmentationabstractExisting methods for video object segmentation (VOS) have achieved significant success by performing semantic guidance, spatial constraint, or temporal consistency. However, VOS still remains highly challenging because it is difficult to collaboratively leverage spatial constraint, temporal consistency, and semantic guidance while reducing redundant information. In this paper, we propose an efficient unified spatio-temporal dynamic routing (STDR) framework to address VOS by achieving a better spatio-temporal balance while avoiding redundancy. Specifically, our unified spatio-temporal modeling contains three paths: 1) short-term spatial path is employed to mine the spatial constraints from the previous frame; 2) long-term semantic path is used to capture semantic cues from the first reference frame with ground-truth labels; 3) memory queue path is designed to efficiently exploit the temporal consistency of middle frames with a compact memory bank of constant size. To enhance the input of each path, we introduce a progressive contextual memory enhancement module to exploit the contextualized memory with growing receptive fields by progressively aggregating spatial contextual information from adjacent frames for each memory frame. Furthermore, we design a dynamic memory-routed module to globally refine the outputs of our three paths for unified modeling. Enhanced by the proposed modules, our STDR achieves state-of-the-art performance with fast speed on the DAVIS 2016, DAVIS 2017 Val/Test, YouTube-VOS 2018/2019, and real-world long-video benchmarks. Jisheng Dang, Huicheng Zheng, Xiaohao Xu, Yulan Guo |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Inserting Anybody in Diffusion Models via Celeb BasisabstractExquisite demand exists for customizing the pretrained large text-to-image model, $e.g.$ Stable Diffusion, to generate innovative concepts, such as the users themselves. However, the newly-added concept from previous customization methods often shows weaker combination abilities than the original ones even given several images during training. We thus propose a new personalization method that allows for the seamless integration of a unique individual into the pre-trained diffusion model using just $one\ facial\ photograph$ and only $1024\ learnable\ parameters$ under $3\ minutes$. So we can effortlessly generate stunning images of this person in any pose or position, interacting with anyone and doing anything imaginable from text prompts. To achieve this, we first analyze and build a well-defined celeb basis from the embedding space of the pre-trained large text encoder. Then, given one facial photo as the target identity, we generate its own embedding by optimizing the weight of this basis and locking all other parameters. Empowered by the proposed celeb basis, the new identity in our customized model showcases a better concept combination ability than previous personalization methods. Besides, our model can also learn several new identities at once and interact with each other where the previous customization model fails to. Project page is at: http://celeb-basis.github.io. Code is at: https://github.com/ygtxr1997/CelebBasis. Ge Yuan, Xiaodong Cun, Yong Zhang 0034, Maomao Li, Xintao Wang 0002, Ying Shan, Huicheng Zheng |
NeurIPS | 8 |
| 2023 | Dual-Memory Feature Aggregation for Video Object Detection
Diwei Fan, Huicheng Zheng, Jisheng Dang |
PRCV (6) | 2 |
| 2023 | Multiple Histograms-Based Reversible Data Hiding Using Fast Performance Optimization and Adaptive Pixel DistributionabstractAbstract In prediction error-based reversible data hiding, multiple histograms modification (MHM) is well known for high image quality and thus has received wide attention in recent years. However, the computational cost for performance optimization in MHM is too high, which is particularly critical for real-time applications. This manuscript aims to reduce the computational complexity of MHM by presenting two techniques, including fast performance optimization and adaptive pixel distribution. Fast performance optimization provides a two-stage process for optimal bin selection by exploiting the concept of per-bit distortion of data embedding within a prediction error histogram (PEH). In fast performance optimization, the distribution characteristics of the per-bit distortion are investigated to significantly narrow down the solution space of optimal bin selection. The second technique is adaptive pixel distribution, which tries to nonuniformly allocate pixels into multiple PEHs to further reduce the time complexity. Extensive experiments show that the computational complexity of MHM is significantly reduced while well preserving the image quality. Junying Yuan, Huicheng Zheng, Jiangqun Ni |
Comput. J. | 2 |
| 2023 | Efficient and Robust Video Object Segmentation Through Isogenous Memory Sampling and Frame Relation MiningabstractRecently, memory-based methods have achieved remarkable progress in video object segmentation. However, the segmentation performance is still limited by error accumulation and redundant memory, primarily because of 1) the semantic gap caused by similarity matching and memory reading via heterogeneous key-value encoding; 2) the continuously growing and inaccurate memory through directly storing unreliable predictions of all previous frames. To address these issues, we propose an efficient, effective, and robust segmentation method based on Isogenous Memory Sampling and Frame-Relation mining (IMSFR). Specifically, by utilizing an isogenous memory sampling module, IMSFR consistently conducts memory matching and reading between sampled historical frames and the current frame in an isogenous space, minimizing the semantic gap while speeding up the model through an efficient random sampling. Furthermore, to avoid key information loss during the sampling process, we further design a frame-relation temporal memory module to mine inter-frame relations, thereby effectively preserving contextual information from the video sequence and alleviating error accumulation. Extensive experiments demonstrate the effectiveness and efficiency of the proposed IMSFR method. In particular, our IMSFR achieves state-of-the-art performance on six commonly used benchmarks in terms of region similarity & contour accuracy and speed. Our model also exhibits strong robustness against frame sampling due to its large receptive field. Jisheng Dang, Huicheng Zheng, Jinming Lai, Xu Yan 0005, Yulan Guo |
IEEE Trans. Image Process. | 2 |
| 2022 | MSML: Enhancing Occlusion-Robustness by Multi-Scale Segmentation-Based Mask Learning for Face RecognitionabstractIn unconstrained scenarios, face recognition remains challenging, particularly when faces are occluded. Existing methods generalize poorly due to the distribution distortion induced by unpredictable occlusions. To tackle this problem, we propose a hierarchical segmentation-based mask learning strategy for face recognition, enhancing occlusion-robustness by integrating segmentation representations of occlusion into face recognition in the latent space. We present a novel multi-scale segmentation-based mask learning (MSML) network, which consists of a face recognition branch (FRB), an occlusion segmentation branch (OSB), and hierarchical elaborate feature masking (FM) operators. With the guidance of hierarchical segmentation representations of occlusion learned by the OSB, the FM operators can generate multi-scale latent masks to eliminate mistaken responses introduced by occlusions and purify the contaminated facial features at multiple layers. In this way, the proposed MSML network can effectively identify and remove the occlusions from feature representations at multiple levels and aggregate features from visible facial areas. Experiments on face verification and recognition under synthetic or realistic occlusions demonstrate the effectiveness of our method compared to state-of-the-art methods. Ge Yuan, Huicheng Zheng, Jiayu Dong |
AAAI | 2 |
| 2022 | Variance of Local Contribution: an Unsupervised Image Quality Assessment for Face RecognitionabstractIn recent years, Face Image Quality Assessment (FIQA) plays an important role in the face recognition system. However, how to define face image quality is still an open question. In this work, we argue that a high-quality face image should have more identity-related information than a low-quality face image. Thus, we propose a novel unsupervised Face Image Quality Assessment with the variance of local contribution (VLC-FIQA). In our approach, we alternately mask partial pixels of the face image, then quantify the importance of these pixels and compute the variation of the importance of different parts as the quality of the image. Extensive experiments show that our VLC-FIQA outperforms state-of-the-art approaches on LFW. Our approach can be easily used for any recognition system and be extended to other recognition tasks such as person re-identification. Qiye Lian, Xiaohua Xie, Huicheng Zheng, Yongdong Zhang 0002 |
ICPR | 3 |
| 2022 | Detail injection with heterogeneous composite backbone network for object detection
Zhiwei Yan, Huicheng Zheng |
Multim. Tools Appl. | 2 |
| 2021 | From Coarse to Fine: Hierarchical Multi-scale Temporal Information Modeling via Sub-group Convolution for Video Action RecognitionabstractIn the video action recognition task, it is essential to model the temporal information. Since different actions have different durations, capturing multi-scale temporal features is very crucial. In this paper, we propose a multi-scale modeling (MSM) module to exploit temporal information for action recognition, which is composed of a multi-scale temporal convolution (MTC) block and a multi-scale hierarchical convolution (MHC) block. MTC uses convolutions of multiple temporal depths to capture features of different temporal scales to enhance the connection between frames. MHC employs group convolution to obtain more fine-grained multi-scale features in the channel dimension. In MHC, the convolutions are formulated as a hierarchical structure to expand the receptive fields as well as help implement a series of sub-group convolutions, which can help to realize the modeling of distinctive and long-term information. The two components of MSM are complementary in temporal modeling. Finally, we evaluated our method on several action recognition benchmarks including Kinetics, UCF10l, and HMDB51, and obtained competitive results, which verified the effectiveness of the proposed method in temporal modeling, Fengwen Cheng, Huicheng Zheng, Zehua Liu |
IJCNN | 2 |
| 2021 | Dense Video Captioning with Hierarchical Attention-Based Encoder-Decoder NetworksabstractDense video captioning is a challenging task with the goal of localizing and describing all events in an untrimmed video, taking into account both visual and text information. Although existing methods have made some achievements, most of them suffer from missing details and inferior captioning. Recent progress has been made in using object features to supplement more detailed information. However, due to the considerable number of objects in the video, the representation of learning objects is often noisy, which may interfere with the generation of correct captions. We also notice that realworld video-text data involve different granularity levels, such as objects/words and events/sentences. Therefore, we propose the hierarchical video-text attention-based encoder-decoder networks for dense video captioning. The proposed method successfully considers the hierarchy in the video and text and exploits the most relevant visual and text features when generating caption. Specially, we design a hierarchical attention encoder for learning complex visual information: an object attention module focusing on the most relevant objects and an event attention module modeling the long-range temporal context. A corresponding decoder has been built for translating multi-level features into the linguistic description, i.e., a word attention module to exploit the most correlated text features and a sentence attention module to leverage high-level semantic information. The proposed hierarchical attention mechanism achieves state-of-the-art performance on the ActivityNet Captions dataset. Mingjing Yu, Huicheng Zheng, Zehua Liu |
IJCNN | 2 |
| 2021 | Foreground Feature Selection and Alignment for Adaptive Object Detection
Huicheng Zheng, Manwei Chen |
PRCV (1) | 2 |
| 2021 | A Residual Correction Approach for Semi-supervised Semantic Segmentation
Haoliang Li, Huicheng Zheng |
PRCV (4) | 2 |
| 2021 | Detection-Oriented Backbone Trained from Near Scratch and Local Feature Refinement for Small Object Detection
Zhiwei Yan, Huicheng Zheng, Lvran Chen |
Neural Process. Lett. | 2 |
| 2021 | Image stitching based on angle-consistent warping
Yinqi Chen, Huicheng Zheng, Yiyan Ma, Zhiwei Yan |
Pattern Recognit. | 2 |
| 2021 | Event-Centric Hierarchical Representation for Dense Video CaptioningabstractDense video captioning aims to localize and describe multiple events in untrimmed videos, which is a challenging task that draws attention recently in computer vision. Although existing methods have achieved impressive performance, most of them only focus on local information of event segments or very simple event-level context, overlooking the complexity of event-event relationship and the holistic scene. As a result, the coherence of captions within the same video could be damaged. In this article, we propose a novel event-centric hierarchical representation to alleviate this problem. We enhance the event-level representation by capturing rich relationship between events in terms of both temporal structure and semantic meaning. Then, a caption generator with late fusion is developed to generate surrounding-event-aware and topic-aware sentences, conditioned on the hierarchical representation of visual cues from the scene level, the event level, and the frame level. Furthermore, we propose a duplicate removal method, namely temporal-linguistic non-maximum suppression (TL-NMS) to distinguish redundancy in both localization and captioning stages. Quantitative and qualitative evaluations on the ActivityNet Captions and YouCook2 datasets demonstrate that our method improves the quality of generated captions and achieves state-of-the-art performance on most metrics. Teng Wang 0007, Huicheng Zheng, Mingjing Yu, Qian Tian, Haifeng Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Discriminative Region Mining for Object DetectionabstractIn generic object detection, detectors are often susceptible to foreground objects and background regions that share similar appearances. In this paper, we propose a novel discriminative region mining (DRM) module for object detection, which enables discriminative region localization and representation for accurate object identification. The DRM module is collaboratively optimized by an extra intramodule classification loss in addition to the usual detection loss, which ensures its adequate discriminative capability. Specifically, two derivatives of the DRM module, namely a local DRM module and a contextual DRM module are proposed to excavate local and contextual discriminative regions, respectively. Furthermore, we extend the local DRM module to capture multiple local discriminative regions with a diversity constraint. To explore informative local features, an image upsampling branch is introduced to generate fine-grained representation for the local DRM module. Extensive experiments on the PASCAL VOC and MS COCO datasets demonstrate the effectiveness of the proposed method. Simple baseline detectors with the built-in DRM can achieve state-of-the-art detection performance. For example, the proposed detector achieves a mean average precision of 81.0% on PASCAL VOC 2007 with an input size of$\text{300} \times \text{300}$using a ResNet-18 backbone, which runs at 24.2 fps on an Nvidia Titan X GPU. Lvran Chen, Huicheng Zheng, Zhiwei Yan |
IEEE Trans. Multim. | 2 |
| 2020 | Feature Enhancement for Multi-scale Object Detection
Huicheng Zheng, Lvran Chen, Zhiwei Yan |
Neural Process. Lett. | 1 |
| 2020 | Facial expression recognition based on a multi-task global-local network
Mingjing Yu, Huicheng Zheng, Zhifeng Peng, Jiayu Dong, Heran Du |
Pattern Recognit. Lett. | 2 |
| 2020 | Laplacian-Uniform Mixture-Driven Iterative Robust Coding With Applications to Face Recognition Against Dense ErrorsabstractOutliers due to occlusion, pixel corruption, and so on pose serious challenges to face recognition despite the recent progress brought by sparse representation. In this article, we show that robust statistics implemented by the state-of-the-art methods are insufficient for robustness against dense gross errors. By modeling the distribution of coding residuals with a Laplacian-uniform mixture, we obtain a sparse representation that is significantly more robust than the previous methods. The nonconvex error term of the implemented objective function is nondifferentiable at zero and cannot be properly addressed by the usual iteratively reweighted least-squares formulation. We show that an iterative robust coding algorithm can be derived by local linear approximation of the nonconvex error term, which is both effective and efficient. With iteratively reweighted l1minimization of the error term, the proposed algorithm is capable of handling the sparsity assumption of the coding errors more appropriately than the previous methods. Notably, it has the distinct property of addressing error detection and error correction cooperatively in the robust coding process. The proposed method demonstrates significantly improved robustness for face recognition against dense gross errors, either contiguous or discontiguous, as verified by extensive experiments. Huicheng Zheng, Dajun Lin, Lina Lian, Jiayu Dong |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Low-Rank Laplacian-Uniform Mixed Model for Robust Face RecognitionabstractSparse representation based methods have successfully put forward a general framework for robust face recognition through linear reconstruction and sparsity constraints. However, residual modeling in existing works is not yet robust enough when dealing with dense noise. In this paper, we aim at recognizing identities from faces with varying levels of noises of various forms such as occlusion, pixel corruption, or disguise, and take improving the fitting ability of the error model as the key to addressing this problem. To fully capture the characteristics of different noises, we propose a mixed model combining robust sparsity constraint and low-rank constraint, which can deal with random errors and structured errors simultaneously. For random noises such as pixel corruption, we adopt a Laplacian-uniform mixed function for fitting the error distribution. For structured errors like continuous occlusion or disguise, we utilize robust nuclear norm to constrain the rank of the error matrix. An effective iterative reweighted algorithm is then developed to solve the proposed model. Comprehensive experiments were conducted on several benchmark databases for robust face recognition, and the overall results demonstrate that our model is most robust against various kinds of noises, when compared with state-of-the-art methods. Jiayu Dong, Huicheng Zheng, Lina Lian |
CVPR | 2 |
| 2019 | Detail preservation and feature refinement for object detection
Huicheng Zheng, Zhiwei Yan, Lvran Chen |
Neurocomputing | 2 |
| 2019 | Multi-object Tracking by Joint Detection and Identification Learning
Bo Ke, Huicheng Zheng, Lvran Chen, Zhiwei Yan |
Neural Process. Lett. | 2 |
| 2019 | Cross-Line Pedestrian Counting Based on Spatially-Consistent Two-Stage Local Crowd Density Estimation and AccumulationabstractThis paper proposes a scalable approach for counting pedestrians crossing a virtual line when the crowd is highly dynamic and possibly extremely dense. The approach mainly consists of two parts: local crowd density estimation and pedestrian counting based on accumulating local densities across the line. To obtain a fine estimation of local crowd densities, we divide the neighborhood at the line into a number of blocks. We enforce spatial consistency between local counts in the blocks and those in the enclosing regions to guarantee consistent estimation of local crowd densities. For scalability to various density levels in crowd density estimation, we propose a two-stage strategy: pre-classification of density levels and subsequent regression with overlapped operational ranges. To count pedestrians crossing the virtual line, we accumulate the crowd densities across the line according to the locally estimated velocities. Extensive experimental results demonstrate the effectiveness of the proposed approach and its scalability to crowdedness. Huicheng Zheng, Zijian Lin, Jiepeng Cen, Yadan Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Dynamic Facial Expression Recognition Based on Convolutional Neural Networks with Dense ConnectionsabstractFacial expression recognition (FER) is a challenging problem with important applications. Applying deep learning techniques for dynamic FER is advantageous in terms of automatically generating discriminative expression features. However, inadequate training data in small expression databases aggravates overfitting and hinders the performance of deep networks. Recent studies have shown that dense connectivity in convolutional neural networks can encourage feature sharing and alleviate overfitting when training with small datasets. Still, traditional dense structures are too deep for FER with insufficient training samples. In this paper, we propose a relatively shallow CNN structure with densely connected short paths for FER. Instead of using transition layers to down-sample feature maps between dense blocks, we introduce dense connectivity across pooling to enforce feature sharing in the shallow CNN structure. Extensive experiments show that our method achieves competitive performance on benchmark datasets CK+ and Oulu-CASIA. Jiayu Dong, Huicheng Zheng, Lina Lian |
ICPR | 2 |
| 2018 | Temporal Inception Architecture for Action Recognition with Convolutional Neural NetworksabstractModeling appearance and short-term dynamic information is the mainstream strategy for action recognition based on deep learning. We consider it important to model the multi-scale temporal information, including both short-term information and long-term information, for action representation. In this paper, a novel temporal inception architecture (TIA) is proposed to solve this problem, which is a general structure that can be combined with multi-segment-based frameworks for action recognition. The TIA is composed of multiple spatial-temporal convolutional branches, in which the temporal information of different scales is extracted. Then feature maps of all branches are concatenated as the output of TIA. In our experiments, the TIA is embedded into temporal segment networks (TSN) to construct our temporal segment inception networks (TSIN) for action recognition tasks. Extensive experiments demonstrate that TSIN outperforms TSN and achieves the state-of-the-art performance on HMDB51 and UCF101. Jiepeng Cen, Huicheng Zheng |
ICPR | 3 |
| 2018 | Facial Expression Recognition Based on Region-Wise Attention and Geometry Difference
Heran Du, Huicheng Zheng, Mingjing Yu |
PRCV (3) | 2 |
| 2018 | Domain Attention Model for Domain Generalization in Object Detection
Wei-Xiong He, Huicheng Zheng, Jian-Huang Lai |
PRCV (4) | 2 |
| 2018 | Multi-flow Sub-network and Multiple Connections for Single Shot Detection
Huicheng Zheng, Lvran Chen |
PRCV (2) | 2 |
| 2018 | Multi-level Three-Stream Convolutional Networks for Video-Based Action Recognition
Yijing Lv, Huicheng Zheng |
PRCV (2) | 2 |
| 2017 | Object Detection by Learning Oriented Gradients
Huicheng Zheng, Ziqian Luo |
ICIG (2) | 2 |
| 2017 | Adaptive Patch Quantization for Histogram-Based Visual Tracking
Lvran Chen, Huicheng Zheng, Zijian Lin, Dajun Lin, Bo Ke |
ICIG (2) | 2 |
| 2017 | Activation-Based Weight Significance Criterion for Pruning Deep Neural Networks
Jiayu Dong, Huicheng Zheng, Lina Lian |
ICIG (2) | 2 |
| 2017 | Robust face recognition based on iterative sparse coding and pixel selectionabstractFace recognition based on sparse representation has attracted broad interest in recent years. In many existing sparse coding works, the distribution of error term (coding residual) is modeled with a Laplacian or Gaussian function, which leads to an l1-norm or l2-norm minimization problem. However, it is hard to fit the error term satisfactorily in practice, especially when occlusion or corruption exists. In order to improve the robustness of sparse coding algorithms to outliers, we propose an efficient pixel-selection strategy in this paper, which can pick out unspoiled pixels from a seriously damaged face image. However, it is very challenging to determine the positions of the outliers directly. We propose an iterative coding approach to improve the selection. Extensive experiments demonstrate the robustness and effectiveness of the proposed method in the presence of occlusion and corruption. Lina Lian, Huicheng Zheng, Jiayu Dong |
ICIP | 2 |
| 2017 | Online multi-object tracking based on hierarchical association and sparse representationabstractIn recent years, sparse representation has been applied to multi-object tracking and shows promising performance. But existing methods often lead to considerable computation. In this paper, we propose a two-level hierarchical association approach to improve the accuracy and efficiency of online multi-object tracker based on sparse representation. We employ a time-saving affinity measure and a discriminative sparse representation classifier to handle objects with disparate and similar appearances, respectively. We also propose a novel strategy for track termination to protect the reliable tracks containing more detections and restrain the unreliable tracks at the same time. Experimental results demonstrate that the proposed method outperforms state-of-the-art online methods. Zijian Lin, Huicheng Zheng, Bo Ke, Lvran Chen |
ICIP | 2 |
| 2016 | Cross-View Action Recognition Based on a Statistical Translation FrameworkabstractActions captured under view changes pose serious challenges to modern action recognition methods. In this paper, we propose an effective approach for cross-view action recognition based on a statistical translation framework, which boils down to estimation of visual word transfer probabilities across views. Specifically, local features are extracted from action video frames and form bags of words based on k-means clustering. Though the appearance of an action may vary due to view changes, the underlying transfer tendency between visual words across views can be exploited. We propose two methods to measure the visual-word-based transfer relationship that are eventually based on frequency counts of word pairs. In the first method, word transfer probabilities are estimated by maximizing the likelihood of a shared action set with the EM algorithm. In the second method, word transfer probabilities are estimated by using likelihood-ratio tests. The two methods achieve comparable results and perform better when they are combined. For cross-view action classification, we compute action transfer probabilities based on the estimated word transfer probabilities and then implement a K-NN-like classification based on action video transfer probabilities. We verified our method on the public multiview IXMAS dataset and the WVU dataset. Promising results are obtained compared with state-of-the-art methods. Huicheng Zheng, Jinyu Gao, Jiepeng Cen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Cross-view Action Recognition via Dual-Codebook and Hierarchical Transfer Framework
Huicheng Zheng, Jian-Huang Lai |
ACCV (5) | 2 |
| 2013 | Pedestrian Counting Based on Crowd Density Estimation and Lucas-Kanade Optical FlowabstractThis paper proposes a novel method to estimate crowd densities in regions and count pedestrians passing through a virtual line. Firstly, the crowd density estimation for regions is based on regional feature analysis and support vector regression (SVR). We extract the following features from each segmented region: the pixel ratio and block-size histogram of the foreground, the pixel ratio and the Minkowski dimension of the edge image, and the gray-level co-occurrence matrix (GLCM) features of the gray image. SVR is used to train and estimate the crowd densities of regions with the extracted features. Secondly, we use the estimated crowd densities and the Lucas-Kanade (LK) optical flow to count pedestrians passing through the line. In each region, we divide the estimated crowd density by the number of foreground pixels and get the density per foreground pixel. Then, people moving speeds based on the LK optical flow are used to compute the number of foreground pixels crossing the line. After that, we can estimate the number of people passing through the line by multiplying the density per foreground pixel by the number of foreground pixels crossing the line. Each region is segmented into several small blocks for higher accuracy to compute the people moving speeds. The experimental results show that the proposed method has a good performance on both crowd density estimation and pedestrian counting. Huicheng Zheng |
ICIG | 2 |
| 2013 | Person Re-identification by Multi-resolution Saliency-Weighted Color Histograms and Local Structural Sparse CodingabstractPerson re-identification plays an important role in computer vision, aiming to identify the same person viewed by disjoint cameras at different time instants and locations. In this paper we present a novel appearance-based method by multi-resolution saliency-weighted color histograms and local structural sparse coding for re-identification work. The former descriptor captures global chromatic content while the latter exploits both partial and spatial information of individuals. Specifically, visual saliency is considered as weighting operators to increase the discriminative power of features. Finally a combinational matching strategy is employed to measure the similarity between individuals. Experimental results over two challenging benchmark datasets (VIPeR, ETHZ) demonstrate that our method obtains competitive performance. Dandan Xu, Huicheng Zheng |
ICIG | 2 |
| 2013 | Different ZFs Leading to Various ZNN Models Illustrated via Online Solution of Time-Varying Underdetermined Systems of Linear Equations with Robotic Application
Yunong Zhang, Ying Wang 0031, Long Jin 0001, Bingguo Mu, Huicheng Zheng |
ISNN (2) | 5 |
| 2013 | Link Between and Comparison and Combination of Zhang Neural Network and Quasi-Newton BFGS Method for Time-Varying Quadratic MinimizationabstractSince 2001, a novel type of recurrent neural network called Zhang neural network (ZNN) has been proposed, investigated, and exploited for solving online time-varying problems in a variety of scientific and engineering fields. In this paper, three discrete-time ZNN models are first proposed to solve the problem of time-varying quadratic minimization (TVQM). Such discrete-time ZNN models exploit methodologically the time derivatives of time-varying coefficients and the inverse of the time-varying coefficient matrix. To eliminate explicit matrix-inversion operation, the quasi-Newton BFGS method is introduced, which approximates effectively the inverse of the Hessian matrix; thus, three discrete-time ZNN models combined with the quasi-Newton BFGS method (named ZNN-BFGS) are proposed and investigated for TVQM. In addition, according to the criterion of whether the time-derivative information of time-varying coefficients is explicitly known/used or not, these proposed discrete-time models are classified into three categories: 1) models with time-derivative information known (i.e., ZNN-K and ZNN-BFGS-K models), 2) models with time-derivative information unknown (i.e., ZNN-U and ZNN-BFGS-U models), and 3) simplified models without using time-derivative information (i.e., ZNN-S and ZNN-BFGS-S models). The well-known gradient-based neural network is also developed to handle TVQM for comparison with the proposed ZNN and ZNN-BFGS models. Illustrative examples are provided and analyzed to substantiate the efficacy of these proposed models for TVQM. Yunong Zhang, Bingguo Mu, Huicheng Zheng |
IEEE Trans. Cybern. | 3 |
| 2010 | Invariant Feature Set Generation with the Linear Manifold Self-organizing Map
Huicheng Zheng |
ACCV (4) | 1 |
| 2009 | Learning nonlinear manifolds based on mixtures of localized linear manifolds under a self-organizing framework
Huicheng Zheng, Qionghai Dai, Sanqing Hu, Zheming Lu 0001 |
Neurocomputing | 1 |
| 2008 | Locally Linear Online Mapping for Mining Low-Dimensional Data Manifolds
Huicheng Zheng, Qionghai Dai, Sanqing Hu |
PAKDD | 1 |
| 2008 | Fast-Learning Adaptive-Subspace Self-Organizing Map: An Application to Saliency-Based Invariant Image Feature ConstructionabstractThe adaptive-subspace self-organizing map (ASSOM) is useful for invariant feature generation and visualization. However, the learning procedure of the ASSOM is slow. In this paper, two fast implementations of the ASSOM are proposed to boost ASSOM learning based on insightful discussions of the basis rotation operator of ASSOM. We investigate the objective function approximately maximized by the classical rotation operator. We then explore a sequence of two schemes to apply the proposed ASSOM implementations to saliency-based invariant feature construction for image classification. In the first scheme, a cumulative activity map computed from a single ASSOM is used as descriptor of the input image. In the second scheme, we use one ASSOM for each image category and a joint cumulative activity map is calculated as the descriptor. Both schemes are evaluated on a subset of the Corel photo database with ten classes. The multi-ASSOM scheme is favored. It is also applied to adult image filtering and shows promising results. Huicheng Zheng, Grégoire Lefebvre, Christophe Laurent |
IEEE Trans. Neural Networks | 1 |
| 2006 | On the Basis Updating Rule of Adaptive-Subspace Self-Organizing Map (ASSOM)
Huicheng Zheng, Christophe Laurent, Grégoire Lefebvre |
ICANN (1) | 1 |
| 2005 | Skin detection using pairwise models
Bruno Jedynak, Huicheng Zheng, Mohamed Daoudi |
Image Vis. Comput. | 2 |
| 2004 | From Maximum Entropy to Belief Propagation: An application to Skin DetectionabstractWe build a maximum entropy model for skin detection. This model imposes constraints on various marginal distributions. Parameter estimation as well as optimization cannot be tackled without approximations. We propose to use a tree approximation of the pixel lattice. Parameter estimation is then reduced to the estimations of color histograms for neighbor pixels. Moreover, the belief propagation algorithm permits to obtain fast solution for skin probability at pixel locations. We assess the performance on the Compaq database. 1 Huicheng Zheng, Mohamed Daoudi, Bruno Jedynak |
BMVC | 1 |
| 2004 | Blocking objectionable images: adult images and harmful symbolsabstractThis paper describes a practical objectionable image filtering system, aimed at children's safer Web access. It includes two image filters: adult image filter and harmful symbol filter. In the adult image filter, we adopt a statistical model for skin detection and a neural network for adult image classification. The performance of the skin detection of our model outperforms that of the baseline model. Its elapsed time is about 0.18 second per image, which compares very well against previous systems. In the harmful symbol filter, we present an edge based Zernike moments method, which can capture the shape feature of a symbol object effectively. Its elapsed time is about 0.13 second per image. Experimental results on a large image database show that both of our filters can give promising performances Huicheng Zheng, Mohamed Daoudi |
ICME | 1 |