Rui Huang 0001

dblp:56/2875-1 · DBLP profile ↗
← Back
85ranked-venue papers
10as first author
38since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 54 · 8 first-author · 26 since 2021Artificial intelligence and machine learning · 43 · 6 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 3 since 2021Systems, architecture and hardware · 6 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A whole-life fatigue crack growth rate prediction method based on active learning and physics-informed loss
Qixuan Zhang, Wei Zhang 0021, Rui Huang 0001, Xinghui Chen, Changyu Zhou
Eng. Appl. Artif. Intell.3
2026 Seeking Flat Minima Over Diverse Surrogates for Improved Adversarial Transferability: A Theoretical Framework and Algorithmic Instantiation
abstract
The transfer-based black-box adversarial attack setting poses the challenge of crafting an adversarial example (AE) on known surrogate models that remains effective against unseen target models. Due to the practical importance of this task, numerous methods have been proposed to address this challenge. However, most previous methods are heuristically designed and intuitively justified, lacking a theoretical foundation. To bridge this gap, we derive a novel transferability bound that offers provable guarantees for adversarial transferability. Our theoretical analysis has the advantages of (i) deepening our understanding of previous methods by building a general attack framework and (ii) providing guidance for designing an effective attack algorithm. Our theoretical results demonstrate that optimizing AEs toward flat minima over the surrogate model set, while controlling the surrogate-target model shift measured by the adversarial model discrepancy, yields a comprehensive guarantee for AE transferability. The results further lead to a general transfer-based attack framework, within which we observe that previous methods consider only partial factors contributing to the transferability. Algorithmically, inspired by our theoretical results, we first elaborately construct the surrogate model set in which models exhibit diverse adversarial vulnerabilities with respect to AEs to narrow the instantiated adversarial model discrepancy. Then, a model-Diversity-compatible Reverse Adversarial Perturbation (DRAP) is generated to effectively promote the flatness of AEs over diverse surrogate models to improve transferability. Extensive experiments on NIPS2017 and CIFAR-10 datasets against various target models demonstrate the effectiveness of our proposed attack.
Meixi Zheng, Kehan Wu, Yanbo Fan, Rui Huang 0001, Baoyuan Wu
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 CFPNet: Improving Lightweight ToF Depth Completion via Cross-Zone Feature Propagation
abstract
Depth completion using lightweight time-of-fight (ToF) depth sensors is attractive due to their low cost. However, lightweight ToF sensors usually have a limited field of view (FOV) compared with cameras. Thus, only pixels in the zone area of the image can be associated with depth signals. Previous methods fail to propagate depth features from the zone area to the outside-zone area effectively, thus suffering from degraded depth completion performance outside the zone. To this end, this paper proposes the CFPNet to achieve cross-zone feature propagation from the zone area to the outside-zone area with two novel modules. The first is a direct-attention-based propagation module (DAPM), which enforces direct crosszone feature acquisition. The second is a large-kernelbased propagation module (LKPM), which realizes crosszone feature propagation by utilizing convolution layers with kernel sizes up to 31. CFPNet achieves state-of-the-art (SOTA) depth completion performance by combining these two modules properly, as verified by extensive experimental results on the ZJU-L5 dataset. The code is available at https://github.com/denyingmxd/CFPNet.
Laiyan Ding, Hualie Jiang, Rui Huang 0001
3DV4
2025 DEFOM-Stereo: Depth Foundation Model Based Stereo Matching
abstract
Stereo matching is a key technique for metric depth estimation in computer vision and robotics. Real-world challenges like occlusion and non-texture hinder accurate disparity estimation from binocular matching cues. Recently, monocular relative depth estimation has shown remarkable generalization using vision foundation models. Thus, to facilitate robust stereo matching with monocular depth cues, we incorporate a robust monocular relative depth model into the recurrent stereo-matching framework, building a new framework for depth foundation model-based stereo-matching, DEFOM-Stereo. In the feature extraction stage, we construct the combined context and matching feature encoder by integrating features from conventional CNNs and DEFOM. In the update stage, we use the depth predicted by DEFOM to initialize the recurrent disparity and introduce a scale update module to refine the disparity at the correct scale. DEFOM-Stereo is verified to have much stronger zero-shot generalization compared with SOTA methods. Moreover, DEFOM-Stereo achieves top performance on the KITTI 2012, KITTI 2015, Middlebury, and ETH3D benchmarks, ranking 1ston many metrics. In the joint evaluation under the robust vision challenge, our model simultaneously outperforms previous models on the individual benchmarks, further demonstrating its outstanding capabilities.
Hualie Jiang, Zhiqiang Lou, Laiyan Ding, Minglang Tan, Rui Huang 0001
CVPR7
2025 ROA-BEV: 2D Region-Oriented Attention for BEV-based 3D Object Detection
abstract
Vision-based Bird’s-Eye-View (BEV) 3D object detection has recently become popular in autonomous driving. However, objects with a high similarity to the background from a camera perspective cannot be detected well by existing methods. In this paper, we propose a BEV-based 3D Object Detection Network with 2D Region-Oriented Attention (ROA-BEV), which enables the backbone to focus more on feature learning of the regions where objects exist. Moreover, our method further enhances the information feature learning ability of ROA through multi-scale structures. Each block of ROA utilizes a large kernel to ensure that the receptive field is large enough to catch information about large objects. Experiments on nuScenes show that ROA-BEV improves the performance based on BEVDepth. The source codes of this work will be available at https://github.com/DFLyan/ROA-BEV.
Yubao Sun, Laiyan Ding, Rui Huang 0001
IROS4
2025 Self-Supervised Enhancement for Depth from a Lightweight ToF Sensor with Monocular Images
abstract
Depth map enhancement using paired high-resolution RGB images offers a cost-effective solution for improving low-resolution depth data from lightweight ToF sensors. Nevertheless, naively adopting a depth estimation pipeline to fuse the two modalities requires groundtruth depth maps for supervision. To address this, we propose a self-supervised learning framework, SelfToF, which generates detailed and scale-aware depth maps. Starting from an image-based self-supervised depth estimation pipeline, we add low-resolution depth as inputs, design a new depth consistency loss, propose a scale-recovery module, and finally obtain a large performance boost. Furthermore, since the ToF signal sparsity varies in real-world applications, we upgrade SelfToF to SelfToF* with submanifold convolution and guided feature fusion. Consequently, SelfToF* maintain robust performance across varying sparsity levels in ToF data. Overall, our proposed method is both efficient and effective, as verified by extensive experiments on the NYU and ScanNet datasets. The code is available at https://github.com/denyingmxd/selftof.
Laiyan Ding, Hualie Jiang, Rui Huang 0001
IROS4
2025 DSC3D: Deformable Sampling Constraints in Stereo 3D Object Detection for Autonomous Driving
abstract
Camera-based stereo 3D object detection estimates 3D properties of objects with binocular images only, which is a cost-effective solution for autonomous driving. The state-of-the-art methods mainly improve the detection accuracy of general objects by designing ingenious stereo matching algorithms or complex pipeline modules. Moreover, additional fine-grained annotations, such as masks or LiDAR point clouds, are often introduced to deal with the occlusion problems, which brings in high manual costs for this task. To address the detection bottleneck caused by occlusion in a more cost-effective manner, we develop a novel stereo 3D object detection method named DSC3D, which achieves significant improvements for occluded objects without introducing additional supervision. Specifically, we first report the ambiguity in feature sampling, which refers to the presence of noisy features in the sampling for occluded objects. Then, we propose the Epipolar Constraint Deform-Attention (ECDA) module to address the unreliable left-right correspondence computation in stereo matching caused by occlusion, which reweights epipolar features by adaptively aggregating local neighbor information. Furthermore, to ensure that 3D property estimation is based on robust object features, we propose visible regions guided constraint to explicitly guide the offset learning for feature sampling. Extensive experiments conducted on the KITTI benchmark have demonstrated the proposed DSC3D outperforms the state-of-the-art camera-based methods.
Wenzhong Guo, Rui Huang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 ClickAdapter: Integrating Details Into Interactive Segmentation Model With Adapter
abstract
Click-based interactive segmentation is the most concise and widely used data labeling method. While existing interactive segmentation methods excel in handling simple targets, they encounter challenges in obtaining high-quality masks from some complex scenes, even with a large number of clicks. Also, the cost of retraining the model from scratch for special scenarios is unacceptably high. To address these issues, we propose ClickAdapter, a simple yet powerful interactive segmentation model adapter without the need for no pre-training. Through introducing a small number of additional parameters and computations, the adapter module effectively enhanced the ability of interactive segmentation models to obtain high-quality prediction with limited clicks. Specifically, we incorporate a detail extractor that aims to extract spatial correlations and local detail features of images. These fine-grained data are then integrated into a model with our adapter to generate segmentation masks with sharp and precise edges. During the training process, only the parameters of our adapter are learnable, thereby reducing the training cost. Features in special scenarios can also be infused more efficiently. To verify the efficiency and performance advantages of the proposed method, a series of experiments on a wide range of benchmarks were conducted, demonstrating that the proposed algorithm achieved cutting-edge performance compared to current state-of-the-art (SOTA) methods.
Shanghong Li, Yongquan Chen, Rui Huang 0001, Feng Wu 0001, Yingliang Miao
IEEE Trans. Circuits Syst. Video Technol.5
2024 Incremental 3D Reconstruction through a Hybrid Explicit-and-Implicit Representation
abstract
3D reconstruction is an important task in computer vision and is widely used in robotics and autonomous driving. When building large-scale scenes, limitations in computing resources and the difficulty of accessing the entire dataset in a single task are inevitable. Therefore, an incremental reconstruction approach is desired. On the one hand, traditional explicit 3D reconstruction methods such as SLAM and SFM require global optimization, which means that time and space resources increase dramatically with the growth of training data. On the other hand, implicit methods like Neural Radiation Fields (NeRF) suffer from catastrophic forgetting if trained incrementally. In this paper, we incrementally reconstruct 3D models in a hybrid representation, where the density of the radiation field is formulated by a voxel grid, and the view-dependent color information of the points is inferred by a shallow MLP. The expansion of the voxel grid and the distillation of the shallow MLP are efficient in this case. Experimental results demonstrate that our incremental method achieves a level of accuracy on par with approaches employing global optimization techniques.
Panwen Hu, Rui Huang 0001
ICRA4
2024 Towards Cross-View-Consistent Self-Supervised Surround Depth Estimation
abstract
Depth estimation is a cornerstone for autonomous driving, yet acquiring per-pixel depth ground truth for supervised learning is challenging. Self-Supervised Surround Depth Estimation (SSSDE) from consecutive images offers an economical alternative. While previous SSSDE methods have proposed different mechanisms to fuse information across images, few of them explicitly consider the cross-view constraints, leading to inferior performance, particularly in overlapping regions. This paper proposes an efficient and consistent pose estimation design and two loss functions to enhance cross-view consistency for SSSDE. For pose estimation, we propose to use only front-view images to reduce training memory and sustain pose estimation consistency. The first loss function is the dense depth consistency loss, which penalizes the difference between predicted depths in overlapping regions. The second one is the multi-view reconstruction consistency loss, which aims to maintain consistency between reconstruction from spatial and spatial-temporal contexts. Additionally, we introduce a novel flipping augmentation to improve the performance further. Our techniques enable a simple neural model to achieve state-of-the-art performance on the DDAD and nuScenes datasets. Last but not least, our proposed techniques can be easily applied to other methods. The code is available at https://github.com/denyingmxd/CVCDepth.
Laiyan Ding, Hualie Jiang, Jie Li 0098, Yongquan Chen, Rui Huang 0001
IROS5
2024 Divide and Conquer: Improving Multi-Camera 3D Perception With 2D Semantic-Depth Priors and Input-Dependent Queries
abstract
3D perception tasks, such as 3D object detection and Bird's-Eye-View (BEV) segmentation using multi-camera images, have drawn significant attention recently. Despite the fact that accurately estimating both semantic and 3D scene layouts are crucial for this task, existing techniques often neglect the synergistic effects of semantic and depth cues, leading to the occurrence of classification and position estimation errors. Additionally, the input-independent nature of initial queries also limits the learning capacity of Transformer-based models. To tackle these challenges, we propose an input-aware Transformer framework that leverages Semantics and Depth as priors (named SDTR). Our approach involves the use of an S-D Encoder that explicitly models semantic and depth priors, thereby disentangling the learning process of object categorization and position estimation. Moreover, we introduce a Prior-guided Query Builder that incorporates the semantic prior into the initial queries of the Transformer, resulting in more effective input-aware queries. Extensive experiments on the nuScenes and Lyft benchmarks demonstrate the state-of-the-art performance of our method in both 3D object detection and BEV segmentation tasks.
Qingyong Hu, Yongquan Chen, Rui Huang 0001
IEEE Trans. Image Process.5
2024 End-to-End Video Scene Graph Generation With Temporal Propagation Transformer
abstract
Video scene graph generation has been an emerging research topic, which aims to interpret a video as a temporally-evolving graph structure by representing video objects as nodes and their relations as edges. Existing approaches predominantly follow a multi-step scheme, including frame-level object detection, relation recognition and temporal association. Although effective, these approaches neglect the mutual interactions between independent steps, resulting in a sub-optimal solution. We present a novel end-to-end framework for video scene graph generation, which naturally unifies object detection, object tracking, and relation recognition via a new Transformer structure, namely Temporal Propagation Transformer (TPT). Particularly, TPT extends the existing Transformer-based object detector (e.g., DETR) along the temporal dimension by involving a query propagation module, which can additionally associate the detected instances by identities across frames. A temporal dynamics encoder is then leveraged to dynamically enrich the features of the detected instances for relation recognition by attending to their historic states in previous frames. Meanwhile, the relation propagation strategy is devised to emphasize the temporal consistency of relation recognition results among adjacent frames. Extensive experiments conducted on VidHOI and Action Genome benchmarks demonstrate the superior performance of the proposed TPT over the state-of-the-art methods.
Yong Zhang 0056, Yingwei Pan, Ting Yao 0003, Rui Huang 0001, Tao Mei 0001, Chang Wen Chen
IEEE Trans. Multim.4
2024 Denoised Non-Local Neural Network for Semantic Segmentation
abstract
The non-local (NL) network has become a widely used technique for semantic segmentation, which computes an attention map to measure the relationships of each pixel pair. However, most of the current popular NL models tend to ignore the phenomenon that the calculated attention map appears to be very noisy, containing interclass and intraclass inconsistencies, which lowers the accuracy and reliability of the NL methods. In this article, we figuratively denote these inconsistencies as attention noises and explore the solutions to denoise them. Specifically, we inventively propose a denoised NL network, which consists of two primary modules, i.e., the global rectifying (GR) block and the local retention (LR) block, to eliminate the interclass and intraclass noises, respectively. First, GR adopts the class-level predictions to capture a binary map to distinguish whether the selected two pixels belong to the same category. Second, LR captures the ignored local dependencies and further uses them to rectify the unwanted hollows in the attention map. The experimental results on two challenging semantic segmentation datasets demonstrate the superior performance of our model. Without any external training data, our proposed denoised NL can achieve the state-of-the-art performance of 83.5% and 46.69% mean of classwise intersection over union (mIoU) on Cityscapes and ADE20K, respectively.
Jie Li 0098, Rui Huang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 Learning to Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic Space
abstract
Scene graph generation (SGG) aims to abstract an image into a graph structure, by representing objects as graph nodes and their relations as labeled edges. However, two knotty obstacles limit the practicability of current SGG methods in real-world scenarios: 1) training SGG models requires time-consuming ground-truth annotations, and 2) the closed-set object categories make the SGG models limited in their ability to recognize novel objects outside of training corpora. To address these issues, we novelly exploit a powerful pre-trained visual-semantic space (VSS) to trigger language-supervised and open-vocabulary SGG in a simple yet effective manner. Specifically, cheap scene graph supervision data can be easily obtained by parsing image language descriptions into semantic graphs. Next, the noun phrases on such semantic graphs are directly grounded over image regions through region-word alignment in the pre-trained VSS. In this way, we enable open-vocabulary object detection by performing object category name grounding with a text prompt in this VSS. On the basis of visually-grounded objects, the relation representations are naturally built for relation recognition, pursuing open-vocabulary SGG. We validate our proposed approach with extensive experiments on the Visual Genome benchmark across various SGG scenarios (i.e., supervised / language-supervised, closed-set / open-vocabulary). Consistent superior performances are achieved compared with existing methods, demonstrating the potential of exploiting pre-trained VSS for SGG in more practical scenarios.
Yong Zhang 0056, Yingwei Pan, Ting Yao 0003, Rui Huang 0001, Tao Mei 0001, Chang Wen Chen
CVPR4
2023 Synthesizing a Large Scene with Multiple NeRFs
Shenglong Ye, Rui Huang 0001
ICIG (5)3
2023 A Reinforcement Learning-Based Automatic Video Editing Method Using Pre-trained Vision-Language Model
abstract
In this era of videos, automatic video editing techniques attract more and more attention from industry and academia since they can reduce workloads and lower the requirements for human editors. Existing automatic editing systems are mainly scene-or event-specific, e.g., soccer game broadcasting, yet the automatic systems for general editing, e.g., movie or vlog editing which covers various scenes and events, were rarely studied before, and converting the event-driven editing method to a general scene is nontrivial. In this paper, we propose a two-stage scheme for general editing. Firstly, unlike previous works that extract scene-specific features, we leverage the pre-trained Vision-Language Model (VLM) to extract the editing-relevant representations as editing context. Moreover, to close the gap between the professional-looking videos and the automatic productions generated with simple guidelines, we propose a Reinforcement Learning (RL)-based editing framework to formulate the editing problem and train the virtual editor to make better sequential editing decisions. Finally, we evaluate the proposed method on a more general editing task with a real movie dataset. Experimental results demonstrate the effectiveness and benefits of the proposed context representation and the learning ability of our RL-based editing framework.
Panwen Hu, Yongquan Chen, Rui Huang 0001
ACM Multimedia5
2023 Towards Balanced RGB-TSDF Fusion for Consistent Semantic Scene Completion by 3D RGB Feature Completion and a Classwise Entropy Loss Function
Laiyan Ding, Panwen Hu, Jie Li 0098, Rui Huang 0001
PRCV (2)4
2023 Safe semi-supervised clustering based on Dempster-Shafer evidence theory
Haitao Gan, Zhi Yang 0006, Ran Zhou 0002, Zhiwei Ye, Rui Huang 0001
Eng. Appl. Artif. Intell.6
2023 From Front to Rear: 3D Semantic Scene Completion Through Planar Convolution and Attention-Based Network
abstract
Semantic Scene Completion (SSC) aims to reconstruct complete 3D scenes with precise voxel-wise semantics from the single-view incomplete input data, a crucial but highly challenging problem for scene understanding. Although SSC has seen significant progress due to the introduction of 2D semantic priors in recent years, the occluded parts, especially the rear-view of the scenes, are still poorly completed and segmented. To ameliorate this issue, we propose a novel deep learning framework for 3D SSC, named Planar Convolution and Attention-based Network (PCANet), to effectively extend high-precision predictions of the front-view surface to the rear-view occluded areas. Specifically, we decompose the traditional convolutional layer into three successive planar convolutions to form a Planar Convolution Residual (PCR) block, which maintains the planar features of the 3D scene. Afterward, the Planar Attention Module (PAM) is proposed to capture three different planar attentions and harvest the global context from the front surface to the rear occluded areas to improve the overall accuracy. Extensive experiments on the real NYU and NYUCAD datasets and the synthetic SUNCG-RGBD dataset demonstrate that our proposed framework can generate high-quality SSC results in both front and rear views and outperforms the state-of-the-art approaches trained in an end-to-end manner without additional data.
Jie Li 0098, Xiaohu Yan, Yongquan Chen, Rui Huang 0001
IEEE Trans. Multim.5
2023 Boosting Scene Graph Generation with Visual Relation Saliency
abstract
The scene graph is a symbolic data structure that comprehensively describes the objects and visual relations in a visual scene, while ignoring the inherent perceptual saliency of each visual relation (i.e., relation saliency). However, humans often quickly allocate attention to important/salient visual relations in a scene. To align with such human perception of a scene, we explicitly model the perceptual saliency of visual relation in scene graph by upgrading each graph edge (i.e., visual relation) with an attribute of relation saliency. We present a new design, named as Saliency-guided Message Passing (SMP), that boosts the generation of such scene graph structure with the guidance from the visual relation saliency. Technically, an object interaction encoder is first utilized to strengthen object relation representations by jointly exploiting the appearance, semantic, and spatial relations in between. A branch is further leveraged to estimate the relation saliency of each visual relation by ordinal regression. Next, conditioned on the object and relation features (coupled with the estimated relation saliency), our SMP enhances scene graph generation by performing message passing over the objects and the most salient relations. Extensive experiments on VG-KR and VG150 datasets demonstrate the superiority of SMP for the scene graph generation. Moreover, we empirically validate the compelling generalizability of the learned scene graphs via SMP on downstream tasks like cross-model retrieval and image captioning.
Yong Zhang 0056, Yingwei Pan, Ting Yao 0003, Rui Huang 0001, Tao Mei 0001, Chang Wen Chen
ACM Trans. Multim. Comput. Commun. Appl.4
2022 Salient-to-Broad Transition for Video Person Re-identification
abstract
Due to the limited utilization of temporal relations in video re-id, the frame-level attention regions of mainstream methods are partial and highly similar. To address this problem, we propose a Salient-to-Broad Module (SBM) to enlarge the attention regions gradually. Specifically, in SBM, while the previous frames have focused on the most salient regions, the later frames tend to focus on broader regions. In this way, the additional information in broad regions can supplement salient regions, incurring more powerful video-level representations. To further improve SBM, an Integration-and-Distribution Module (IDM) is introduced to enhance frame-level representations. IDM first integrates features from the entire feature space and then distributes the integrated features to each spatial location. SBM and IDM are mutually beneficial since they enhance the representations from video-level and frame-level, respectively. Extensive experiments on four prevalent benchmarks demonstrate the effectiveness and superiority of our method. The source code is available at https://github.com/baist/SINet.
Shutao Bai, Bingpeng Ma, Hong Chang 0001, Rui Huang 0001, Xilin Chen 0001
CVPR4
2022 Exploring Structure-aware Transformer over Interaction Proposals for Human-Object Interaction Detection
abstract
Recent high-performing Human-Object Interaction (HOI) detection techniques have been highly influenced by Transformer-based object detector (i.e., DETR). Nevertheless, most of them directly map parametric interaction queries into a set of HOI predictions through vanilla Transformer in a one-stage manner. This leaves rich interor intra-interaction structure under-exploited. In this work, we design a novel Transformer-style HOI detector, i.e., Structure-aware Transformer over Interaction Proposals (STIP), for HOI detection. Such design decomposes the process of HOI set prediction into two subsequent phases, i.e., an interaction proposal generation is first performed, and then followed by transforming the non-parametric interaction proposals into HOI predictions via a structure-aware Transformer. The structure-aware Transformer upgrades vanilla Transformer by encoding additionally the holistically semantic structure among interaction proposals as well as the locally spatial structure of human/object within each interaction proposal, so as to strengthen HOI predictions. Extensive experiments conducted on V-COCO and HICO-DET benchmarks have demonstrated the effectiveness of STIP, and superior results are reported when comparing with the state-of-the-art HOI detectors. Source code is available at https://github.com/zyong812/STIP.
Yong Zhang 0056, Yingwei Pan, Ting Yao 0003, Rui Huang 0001, Tao Mei 0001, Chang Wen Chen
CVPR4
2022 Deep Semantic Statistics Matching (D2SM) Denoising Network
Kangfu Mei, Vishal M. Patel, Rui Huang 0001
ECCV (7)3
2022 SANet: Statistic Attention Network for Video-Based Person Re-Identification
abstract
Capturing long-range dependencies during feature extraction is crucial for video-based person re-identification (re-id) since it would help to tackle many challenging problems such as occlusion and dramatic pose variation. Moreover, capturing subtle differences, such as bags and glasses, is indispensable to distinguish similar pedestrians. In this paper, we propose a novel and efficacious Statistic Attention (SA) block which can capture both the long-range dependencies and subtle differences. SA block leverages high-order statistics of feature maps, which contain both long-range and high-order information. By modeling relations with these statistics, SA block can explicitly capture long-range dependencies with less time complexity. In addition, high-order statistics usually concentrate on details of feature maps and can perceive the subtle differences between pedestrians. In this way, SA block is capable of discriminating pedestrians with subtle differences. Furthermore, this lightweight block can be conveniently inserted into existing deep neural networks at any depth to form Statistic Attention Network (SANet). To evaluate its performance, we conduct extensive experiments on two challenging video re-id datasets, showing that our SANet outperforms the state-of-the-art methods. Furthermore, to show the generalizability of SANet, we evaluate it on three image re-id datasets and two more general image classification datasets, including ImageNet. The source code is available athttp://vipl.ict.ac.cn/resources/codes/code/SANet_code.zip.
Shutao Bai, Bingpeng Ma, Hong Chang 0001, Rui Huang 0001, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Circuits Syst. Video Technol.4
2021 PLNet: Plane and Line Priors for Unsupervised Indoor Depth Estimation
abstract
Unsupervised learning of depth from indoor monocular videos is challenging as the artificial environment contains many textureless regions. Fortunately, the indoor scenes are full of specific structures, such as planes and lines, which should help guide unsupervised depth learning. This paper proposes PLNet that leverages the plane and line priors to enhance the depth estimation. We first represent the scene geometry using local planar coefficients and impose the smoothness constraint on the representation. Moreover, we enforce the planar and linear consistency by randomly selecting some sets of points that are probably coplanar or collinear to construct simple and effective consistency losses. To verify the proposed method’s effectiveness, we further propose to evaluate the flatness and straightness of the predicted point cloud on the reliable planar and linear regions. The regularity of these regions indicates quality indoor reconstruction. Experiments on NYU Depth V2 and ScanNet show that PLNet outperforms existing methods. The code is available at https://github.com/HalleyJiang/PLNet.
Hualie Jiang, Laiyan Ding, Junjie Hu 0003, Rui Huang 0001
3DV4
2021 AttaNet: Attention-Augmented Network for Fast and Accurate Scene Parsing
abstract
Two factors have proven to be very important to the performance of semantic segmentation models: global context and multi-level semantics. However, generating features that capture both factors always leads to high computational complexity, which is problematic in real-time scenarios. In this paper, we propose a new model, called Attention-Augmented Network (AttaNet), to capture both global context and multi-level semantics while keeping the efficiency high. AttaNet consists of two primary modules: Strip Attention Module (SAM) and Attention Fusion Module (AFM). Viewing that in challenging images with low segmentation accuracy, there are a significantly larger amount of vertical strip areas than horizontal ones, SAM utilizes a striping operation to reduce the complexity of encoding global context in the vertical direction drastically while keeping most of contextual information, compared to the non-local approaches. Moreover, AFM follows a cross-level aggregation strategy to limit the computation, and adopts an attention strategy to weight the importance of different levels of features at each pixel when fusing them, obtaining an efficient multi-level representation. We have conducted extensive experiments on two semantic segmentation benchmarks, and our network achieves different levels of speed/accuracy trade-offs on Cityscapes, e.g., 71 FPS/79.9% mIoU, 130 FPS/78.5% mIoU, and 180 FPS/70.1% mIoU, and leading performance on ADE20K as well.
Kangfu Mei, Rui Huang 0001
AAAI3
2021 Discriminative Clue Alignment Network for Both Image- and Video-Based Person Re-Identification
Panwen Hu, Rui Huang 0001
BMVC3
2021 BiCnet-TKS: Learning Efficient Spatial-Temporal Representation for Video Person Re-Identification
abstract
In this paper, we present an efficient spatial-temporal representation for video person re-identification (reID). Firstly, we propose a Bilateral Complementary Network (BiCnet) for spatial complementarity modeling. Specifically, BiCnet contains two branches. Detail Branch processes frames at original resolution to preserve the detailed visual clues, and Context Branch with a down-sampling strategy is employed to capture long-range contexts. On each branch, BiCnet appends multiple parallel and diverse attention modules to discover divergent body parts for consecutive frames, so as to obtain an integral characteristic of target identity. Furthermore, a Temporal Kernel Selection (TKS) block is designed to capture short-term as well as long-term temporal relations by an adaptive mode. TKS can be inserted into BiCnet at any depth to construct BiCnet-TKS for spatial-temporal modeling. Experimental results on multiple benchmarks show that BiCnet-TKS outperforms state-of-the-arts with about 50% less computations. The source code is available at https://github.com/blue-blue272/BiCnet-TKS.
Ruibing Hou, Hong Chang 0001, Bingpeng Ma, Rui Huang 0001, Shiguang Shan
CVPR4
2021 Learning Camera Localization via Dense Scene Matching
abstract
Camera localization aims to estimate 6 DoF camera poses from RGB images. Traditional methods detect and match interest points between a query image and a prebuilt 3D model. Recent learning-based approaches encode scene structures into a specific convolutional neural network (CNN) and thus are able to predict dense coordinates from RGB images. However, most of them require re-training or re-adaption for a new scene and have difficulties in handling large-scale scenes due to limited network capacity. We present a new method for scene agnostic camera localization using dense scene matching (DSM), where a cost volume is constructed between a query image and a scene. The cost volume and the corresponding coordinates are processed by a CNN to predict dense coordinates. Camera poses can then be solved by PnP algorithms. In addition, our method can be extended to temporal domain, which leads to extra performance boost during testing time. Our scene-agnostic approach achieves comparable accuracy as the existing scene-specific approaches, such as KFNet, on the 7scenes and Cambridge benchmark. This approach also remarkably outperforms state-of-the-art scene-agnostic dense coordinate regression network SANet. The Code is available at https://github.com/Tangshitao/DenseScene-Matching.
Shitao Tang, Chengzhou Tang, Rui Huang 0001, Siyu Zhu 0001, Ping Tan 0002
CVPR3
2021 Reinforcement Learning Based Automatic Personal Mashup Generation
abstract
Synchronized video editing system, or mashup generation from multiple synchronized videos, has gained much attention due to its high efficiency and low cost in processing videos to convey information. However, few of the existing methods focus on generating the personal mashup over a complete timeline from synchronized surveillance videos, which is increasingly demanded for effectively presenting personal activities without violating the privacy of others. To fill this gap, we develop a Reinforcement Learning (RL)-based personal mashup generation system, which assesses the frame quality at a semantic level and formulates the view selection as an RL problem to improve the efficiency in retrieving the mashups with arbitrary beginnings. Furthermore, we propose a framing objective to perform spatial editing, which enables the views to automatically zoom in and out, so as to present the target people more comprehensively. Both qualitative and quantitative analyses are presented to demonstrate the effectiveness of the proposed frame quality measurements, the RL- based algorithm, and the framing objective.
Panwen Hu, Rui Huang 0001
ICME4
2021 SDAN: Squared Deformable Alignment Network for Learning Misaligned Optical Zoom
abstract
Deep Neural Network (DNN) based super-resolution algorithms have greatly improved the quality of the generated images. However, these algorithms often yield significant artifacts when dealing with real-world super-resolution problems due to the difficulty in learning misaligned optical zoom. In this paper, we introduce a Squared Deformable Alignment Network (SDAN) to address this issue. Our network learns squared per-point offsets for convolutional kernels, and then aligns features in corrected convolutional windows based on the offsets. So the misalignment will be minimized by the extracted aligned features. Different from the per-point off-sets used in the vanilla Deformable Convolutional Network (DCN), our proposed squared offsets not only accelerate the offset learning but also improve the generation quality with fewer parameters. Besides, we further propose an efficient cross packing attention layer to boost the accuracy of the learned offsets. It leverages the packing and unpacking operations to enlarge the receptive field of the offset learning and to enhance the ability of extracting the spatial connection between the low-resolution images and the referenced images. Comprehensive experiments show the superiority of our method over other state-of-the-art methods in both computational efficiency and realistic details. Code is available at https://github.com/MKFMIKU/SDAN
Kangfu Mei, Shenglong Ye, Rui Huang 0001
ICME3
2021 IMENet: Joint 3D Semantic Scene Completion and 2D Semantic Segmentation through Iterative Mutual Enhancement
abstract
3D semantic scene completion and 2D semantic segmentation are two tightly correlated tasks that are both essential for indoor scene understanding, because they predict the same semantic classes, using positively correlated high-level features. Current methods use 2D features extracted from early-fused RGB-D images for 2D segmentation to improve 3D scene completion. We argue that this sequential scheme does not ensure these two tasks fully benefit each other, and present an Iterative Mutual Enhancement Network (IMENet) to solve them jointly, which interactively refines the two tasks at the late prediction stage. Specifically, two refinement modules are developed under a unified framework for the two tasks. The first is a 2D Deformable Context Pyramid (DCP) module, which receives the projection from the current 3D predictions to refine the 2D predictions. In turn, a 3D Deformable Depth Attention (DDA) module is proposed to leverage the reprojected results from 2D predictions to update the coarse 3D predictions. This iterative fusion happens to the stable high-level features of both tasks at a late stage. Extensive experiments on NYU and NYUCAD datasets verify the effectiveness of the proposed iterative late fusion scheme, and our approach outperforms the state of the art on both 3D semantic scene completion and 2D semantic segmentation.
Jie Li 0098, Laiyan Ding, Rui Huang 0001
IJCAI3
2021 Domain Composition and Attention for Unseen-Domain Generalizable Medical Image Segmentation
Ran Gu, Jingyang Zhang, Rui Huang 0001, Wenhui Lei, Guotai Wang, Shaoting Zhang 0001
MICCAI (3)3
2021 Automatic segmentation of organs-at-risk from head-and-neck CT using separable convolutional neural network with hard-region-weighted loss
Wenhui Lei, Haochen Mei, Zhengwentai Sun, Shan Ye, Ran Gu, Huan Wang 0015, Rui Huang 0001, Shichuan Zhang, Shaoting Zhang 0001, Guotai Wang
Neurocomputing7
2021 Dynamic multi-channel metric network for joint pose-aware and identity-invariant facial expression recognition
Yuanyuan Liu 0004, Fang Fang 0008, Yongquan Chen, Rui Huang 0001, Run Wang 0002, Bo Wan 0006
Inf. Sci.5
2021 FocusNetv2: Imbalanced large and small organ segmentation with adversarial shape constraint for head and neck CT images
Yunhe Gao, Rui Huang 0001, Yiwei Yang 0001, Kainan Shao, Changjuan Tao, Yuanyuan Chen 0007, Dimitris N. Metaxas, Hongsheng Li 0001, Ming Chen 0030
Medical Image Anal.2
2021 Cross-Modal Knowledge Adaptation for Language-Based Person Search
abstract
In this paper, we present a method named Cross-Modal Knowledge Adaptation (CMKA) for language-based person search. We argue that the image and text information are not equally important in determining a person's identity. In other words, image carries image-specific information such as lighting condition and background, while text contains more modal agnostic information that is more beneficial to cross-modal matching. Based on this consideration, we propose CMKA to adapt the knowledge of image to the knowledge of text. Specially, text-to-image guidance is obtained at different levels: individuals, lists, and classes. By combining these levels of knowledge adaptation, the image-specific information is suppressed, and the common space of image and text is better constructed. We conduct experiments on the CUHK-PEDES dataset. The experimental results show that the proposed CMKA outperforms the state-of-the-art methods.
Rui Huang 0001, Hong Chang 0001, Chuanqi Tan, Bingpeng Ma
IEEE Trans. Image Process.2
2021 CA-Net: Comprehensive Attention Convolutional Neural Networks for Explainable Medical Image Segmentation
abstract
Accurate medical image segmentation is essential for diagnosis and treatment planning of diseases. Convolutional Neural Networks (CNNs) have achieved state-of-the-art performance for automatic medical image segmentation. However, they are still challenged by complicated conditions where the segmentation target has large variations of position, shape and scale, and existing CNNs have a poor explainability that limits their application to clinical decisions. In this work, we make extensive use of multiple attentions in a CNN architecture and propose a comprehensive attention-based CNN (CA-Net) for more accurate and explainable medical image segmentation that is aware of the most important spatial positions, channels and scales at the same time. In particular, we first propose a joint spatial attention module to make the network focus more on the foreground region. Then, a novel channel attention module is proposed to adaptively recalibrate channel-wise feature responses and highlight the most relevant feature channels. Also, we propose a scale attention module implicitly emphasizing the most salient feature maps among multiple scales so that the CNN is adaptive to the size of an object. Extensive experiments on skin lesion segmentation from ISIC 2018 and multi-class segmentation of fetal MRI found that our proposed CA-Net significantly improved the average segmentation Dice score from 87.77% to 92.08% for skin lesion, 84.79% to 87.08% for the placenta and 93.20% to 95.88% for the fetal brain respectively compared with U-Net. It reduced the model size to around 15 times smaller with close or even better accuracy compared with state-of-the-art DeepLabv3+. In addition, it has a much higher explainability than existing networks by visualizing the attention weight maps. Our code is available at https://github.com/HiLab-git/CA-Net.
Ran Gu, Guotai Wang, Tao Song 0002, Rui Huang 0001, Michael Aertsen, Jan Deprest, Sébastien Ourselin, Tom Vercauteren, Shaoting Zhang 0001
IEEE Trans. Medical Imaging4
2020 Learning Memory Augmented Cascading Network for Compressed Sensing of Images
Yubao Sun, Qingshan Liu 0001, Rui Huang 0001
ECCV (22)4
2020 DiPE: Deeper into Photometric Errors for Unsupervised Learning of Depth and Ego-motion from Monocular Videos
abstract
Unsupervised learning of depth and ego-motion from unlabelled monocular videos has recently drawn great attention, which avoids the use of expensive ground truth in the supervised one. It achieves this by using the photometric errors between the target view and the synthesized views from its adjacent source views as the loss. Despite significant progress, the learning still suffers from occlusion and scene dynamics. This paper shows that carefully manipulating photometric errors can tackle these difficulties better. The primary improvement is achieved by a statistical technique that can mask out the invisible or nonstationary pixels in the photometric error map and thus prevents misleading the networks. With this outlier masking approach, the depth of objects moving in the opposite direction to the camera can be estimated more accurately. To the best of our knowledge, such scenarios have not been seriously considered in the previous works, even though they pose a higher risk in applications like autonomous driving. We also propose an efficient weighted multi-scale scheme to reduce the artifacts in the predicted depth maps. Extensive experiments on the KITTI dataset show the effectiveness of the proposed approaches. The overall system achieves state-of-the-art performance on both depth and ego-motion estimation.
Hualie Jiang, Laiyan Ding, Zhenglong Sun 0001, Rui Huang 0001
IROS4
2020 Multi-organ Segmentation via Co-training Weight-Averaged Models from Few-Organ Datasets
Rui Huang 0001, Yuanjie Zheng, Shaoting Zhang 0001, Hongsheng Li 0001
MICCAI (4)1
2019 How Effectively can Indoor Wireless Positioning Relieve Visual Tracking Pains: A Cramer-Rao Bound Viewpoi
abstract
Visual tracking is fragile in some difficult scenarios, for instance, appearance ambiguity and variation, occlusion can easily degrade most of visual trackers to some extent. In this paper, visual tracking is empowered with wireless positioning to achieve high accuracy while maintaining robustness. Fundamentally different from the previous works, this study does not involve any specific wireless positioning algorithms. Instead, we use the confidence region derived from the wireless positioning Cramér-Rao bound (CRB) as the search region of visual trackers. The proposed framework is low-cost and very simple to implement, yet readily leads to enhanced and robustified visual tracking performance in difficult scenarios as demonstrated by our experimental results. Most importantly, it is utmost valuable for the practioners to pre-evaluate how effectively can the wireless resources available at hand alleviate the visual tracking pains.
Panwen Hu, Zizheng Yan, Rui Huang 0001, Feng Yin 0001
ICIP3
2019 High Quality Monocular Depth Estimation Via A Multi-Scale Network And A Detail-Preserving Objective
abstract
Monocular depth estimation is an important and challenging task in computer vision. Significant progress has been made recently due to deep convolutional neural networks. However, esitmating depth maps with high quality lacks sufficient attention. This paper proposes to recover detailed depth map by training a multi-scale network architecture with a detailpreserving loss function. Firstly, we construct our architecture inspired by the design of atrous spatial pyramid pooling for semantic segmentation. Secondly, we simplify the loss on depth map gradients for preserving details. Experiments on the NYU Depth V2 dataset show that our approach is effective and it achieves state-of-the-art performance, especially in the root mean squared error.
Hualie Jiang, Rui Huang 0001
ICIP2
2019 FocusNet: Imbalanced Large and Small Organ Segmentation with an End-to-End Deep Neural Network for Head and Neck CT Images
Yunhe Gao, Rui Huang 0001, Ming Chen 0030, Zhe Wang 0006, Jincheng Deng, Yuanyuan Chen 0007, Yiwei Yang 0001, Chanjuan Tao, Hongsheng Li 0001
MICCAI (3)2
2019 Confidence-weighted safe semi-supervised clustering
Haitao Gan, Yingle Fan, Zhizeng Luo, Rui Huang 0001, Zhi Yang 0006
Eng. Appl. Artif. Intell.4
2019 A study on multi-kernel intuitionistic fuzzy C-means clustering with multiple attributes
Shan Zeng, Zhiyong Wang 0001, Rui Huang 0001, David Dagan Feng
Neurocomputing3
2018 Multiple Object Tracking by Learning Feature Representation and Distance Metric Jointly
Guoshuai Zhang, Nong Sang, Rui Huang 0001, Jianhua Hou
BMVC4
2018 Active Image-Based Modeling with a Toy Drone
abstract
Image-based modeling techniques [1]-[3] can now generate photo-realistic 3D models from images. But it is up to users to provide high quality images with good coverage and view overlap, which makes the data capturing process tedious and time consuming. We seek to automate data capturing for image-based modeling. The core of our system is an iterative linear method to solve the multi-view stereo (MVS) problem quickly and plan the Next-Best-View (NBV) effectively. Our fast MVS algorithm enables online model reconstruction and quality assessment to determine the NBVs on the fly. We test our system with a toy unmanned aerial vehicle (UAV) in simulated, indoor and outdoor experiments. Results show that our system improves the efficiency of data acquisition and ensures the completeness of the final model.
Rui Huang 0001, Danping Zou, Richard Vaughan 0001, Ping Tan 0002
ICRA1
2018 Safety-aware Graph-based Semi-Supervised Learning
Haitao Gan, Zhizeng Luo, Rui Huang 0001
Expert Syst. Appl.5
2018 On using supervised clustering analysis to improve classification performance
Haitao Gan, Rui Huang 0001, Zhizeng Luo, Xugang Xi, Yunyuan Gao
Inf. Sci.2
2017 Robust Visual Tracking Using Exemplar-Based Detectors
abstract
Tracking by detection has become an attractive tracking technique, which treats tracking as an object detection problem and trains a detector to separate the target object from the background in each frame. While this strategy is effective to some extent, we argue that the task in tracking should be searching for a specific object instance instead of an object category. Based on this viewpoint, a novel framework based on object exemplar detectors is proposed for visual tracking. To build a specific and discriminative model to separate the object instance from the background, the proposed method trains an exemplar-based linear discriminant analysis (ELDA) classifier for the object exemplar, using the current tracked instance as the positive sample and massive negative samples obtained both offline and online. To improve the trackers' adaptivity, we use an ensemble of the above ELDA detectors and update them during the tracking to cover the variation in object appearance. Extensive experimental results on a large benchmark data set show that the proposed method outperforms many state-of-the-art trackers, demonstrating the effectiveness and robustness of the ELDA tracker.
Changxin Gao, Jin-Gang Yu, Rui Huang 0001, Nong Sang
IEEE Trans. Circuits Syst. Video Technol.4
2017 DeepList: Learning Deep Features With Adaptive Listwise Constraint for Person Reidentification
abstract
Person reidentification (re-id) aims to match a specific person across nonoverlapping cameras, which is an important but challenging task in video surveillance. Conventional methods mainly focus either on feature constructing or metric learning. Recently, some deep learning-based methods have been proposed to learn image features and similarity measures jointly. However, current deep models for person re-id are usually trained with eitherpairwise loss, where the number of negative pairs greatly outnumbering that of positive pairs may lead the training model to be biased toward negative pairs orconstant margin hinge loss, without considering the fact that hard negative samples should be paid more attention in the training stage. In this paper, we propose to learn deep representations with an adaptive margin listwise loss. First, ranking lists instead of image pairs are used as training samples, in this way, the problem of data imbalance is relaxed. Second, by introducing an adaptive margin parameter in the listwise loss function, it can assign larger margins to harder negative samples, which can be interpreted as an implementation of the automatic hard negative mining strategy. To gain robustness against changes in poses and part occlusions, our architecture combines four convolutional neural networks, each of which embeds images from different scales or different body parts. The final combined model performs much better than each single model. The experimental results show that our approach achieves very promising results on the challenging CUHK03, CUHK01, and VIPeR data sets.
Jin Wang 0019, Zheng Wang 0007, Changxin Gao, Nong Sang, Rui Huang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2017 Video-Based Pedestrian Re-Identification by Adaptive Spatio-Temporal Appearance Model
abstract
Pedestrian re-identification is a difficult problem due to the large variations in a person's appearance caused by different poses and viewpoints, illumination changes, and occlusions. Spatial alignment is commonly used to address these issues by treating the appearance of different body parts independently. However, a body part can also appear differently during different phases of an action. In this paper, we consider the temporal alignment problem, in addition to the spatial one, and propose a new approach that takes the video of a walking person as input and builds a spatiotemporal appearance representation for pedestrian re-identification. Particularly, given a video sequence, we exploit the periodicity exhibited by a walking person to generate a spatiotemporal body-action model, which consists of a series of body-action units corresponding to certain action primitives of certain body parts. Fisher vectors are learned and extracted from individual body-action units and concatenated into the final representation of the walking person. Unlike previous spatiotemporal features that only take into account local dynamic appearance information, our representation aligns the spatiotemporal appearance of a pedestrian globally. Extensive experiments on public data sets show the effectiveness of our approach compared with the state of the art.
Wei Zhang 0021, Bingpeng Ma, Kan Liu 0001, Rui Huang 0001
IEEE Trans. Image Process.4
2016 Contextual Similarity Regularized Metric Learning for person re-identification
abstract
Person re-identification, aiming to match a specific person among non-overlapping cameras, has attracted plenty of attention in recent years. It can be regarded as a visual retrieval task, namely given a query person image, ranking all gallery images according to their similarities to the query. Conventionally, this similarity function is learnt by forcing intra-distances to be small while inter-distances to be large, which are referred to as individual similarity constraints. In this paper, we propose to learn the similarity function by taking into account of both individual similarity constraints and contextual similarity constraints. The context of a query is defined as its k-nearest neighbors in the gallery. We argue that if two images are from the same person, apart from the visual likeness between them, denoted as the individual similarity, they should also possess similar k-nearest neighbors in the gallery, denoted as the contextual similarity. Motivated by this assumption, we propose a new Contextual Similarity Regularized Metric Learning (CSRML) method for person re-identification. The contextual similarity regularization term forces two images of the same person to share similar context. Both individual and contextual similarity constraints are encoded by a large margin logistic loss function and the final problem is solved by the stochastic gradient descent algorithm. Experiments on the challenging VIPeR and CUHK01 datasets show that our approach achieves very competitive performance.
Jin Wang 0019, Junkang Zhu, Zheng Wang 0007, Changxin Gao, Nong Sang, Rui Huang 0001
ICPR6
2016 Towards designing risk-based safe Laplacian Regularized Least Squares
Haitao Gan, Zhizeng Luo, Xugang Xi, Nong Sang, Rui Huang 0001
Expert Syst. Appl.6
2016 Towards a probabilistic semi-supervised Kernel Minimum Squared Error algorithm
Haitao Gan, Rui Huang 0001, Zhizeng Luo, Yingle Fan, Farong Gao
Neurocomputing2
2016 Image retrieval using spatiograms of colors quantized by Gaussian Mixture Models
Shan Zeng, Rui Huang 0001, Haibing Wang, Zhen Kang
Neurocomputing2
2016 Spatial multi-scale gradient orientation consistency for place instance and Scene category recognition
Changxin Gao, Nong Sang, Rui Huang 0001
Inf. Sci.3
2016 Learning structure of stereoscopic image for no-reference quality assessment with convolutional neural network
Wei Zhang 0021, Chenfei Qu, Lin Ma 0002, Jingwei Guan, Rui Huang 0001
Pattern Recognit.5
2016 Hough Forest-based Association Framework with Occlusion Handling for Multi-Target Tracking
abstract
This letter presents a novel multi-target tracking approach consisting of two parts. The first part is the detection based association to form global tracks. Short yet reliable tracklets are firstly generated. By effectively combining appearance and motion information, a Hough forest learning framework is constructed to obtain a more discriminative affinity model and produce longer association between tracklets. In the second part, in order to connect isolate detections for trajectory consistency, we present an appearance similarity model based on mutual occlusion reasoning. A novel fusion feature template is designed to accurately compute the matching score between each isolated detection and target. Experimental results show significant improvements of our method when compared with several state-of-the-art methods.
Nong Sang, Jianhua Hou, Rui Huang 0001, Changxin Gao
IEEE Signal Process. Lett.4
2016 Multitarget Tracking Using Hough Forest Random Field
abstract
This paper presents a novel tracking-by-detection approach for multitarget tracking. There are two major steps in our framework: data association to form global tracklet association, followed by trajectory estimation to deal with the remaining gaps. In the first step, we formulate tracklet association as an inference problem in a Hough forest random field, which combines Hough forest and conditional random field and allows us to model both local and global tracklet relationships in one unified model. In the second step, we improve the reversible-jump Markov chain Monte Carlo particle filtering method with explicit mutual-occlusion reasoning to fill in the remaining gaps from the first step and increase the overall tracking precision. Extensive experiments have been conducted on five public data sets, and the performance is comparable to that of the state-of-the-art method, if not better.
Nong Sang, Jianhua Hou, Rui Huang 0001, Changxin Gao
IEEE Trans. Circuits Syst. Video Technol.4
2016 Recognizing Focal Liver Lesions in CEUS With Dynamically Trained Latent Structured Models
abstract
This work investigates how to automatically classify Focal Liver Lesions (FLLs) into three specific benign or malignant types in Contrast-Enhanced Ultrasound (CEUS) videos, and aims at providing a computational framework to assist clinicians in FLL diagnosis. The main challenge for this task is that FLLs in CEUS videos often show diverse enhancement patterns at different temporal phases. To handle these diverse patterns, we propose a novel structured model, which detects a number of discriminative Regions of Interest (ROIs) for the FLL and recognize the FLL based on these ROIs. Our model incorporates an ensemble of local classifiers in the attempt to identify different enhancement patterns of ROIs, and in particular, we make the model reconfigurable by introducing switch variables to adaptively select appropriate classifiers during inference. We formulate the model learning as a non-convex optimization problem, and present a principled optimization method to solve it in a dynamic manner: the latent structures (e.g. the selections of local classifiers, and the sizes and locations of ROIs) are iteratively determined along with the parameter learning. Given the updated model parameters in each step, the data-driven inference is also proposed to efficiently determine the latent structures by using the sequential pruning and dynamic programming method. In the experiments, we demonstrate superior performances over the state-of-the-art approaches. We also release hundreds of CEUS FLLs videos used to quantitatively evaluate this work, which to the best of our knowledge forms the largest dataset in the literature. Please find more information at "http://vision.sysu.edu.cn/projects/fllrecog/".
Xiaodan Liang, Liang Lin 0004, Qingxing Cao, Rui Huang 0001, Yongtian Wang
IEEE Trans. Medical Imaging4
2015 A Spatio-Temporal Appearance Representation for Viceo-Based Pedestrian Re-Identification
abstract
Pedestrian re-identification is a difficult problem due to the large variations in a person's appearance caused by different poses and viewpoints, illumination changes, and occlusions. Spatial alignment is commonly used to address these issues by treating the appearance of different body parts independently. However, a body part can also appear differently during different phases of an action. In this paper we consider the temporal alignment problem, in addition to the spatial one, and propose a new approach that takes the video of a walking person as input and builds a spatio-temporal appearance representation for pedestrian re-identification. Particularly, given a video sequence we exploit the periodicity exhibited by a walking person to generate a spatio-temporal body-action model, which consists of a series of body-action units corresponding to certain action primitives of certain body parts. Fisher vectors are learned and extracted from individual body-action units and concatenated into the final representation of the walking person. Unlike previous spatio-temporal features that only take into account local dynamic appearance information, our representation aligns the spatio-temporal appearance of a pedestrian globally. Extensive experiments on public datasets show the effectiveness of our approach compared with the state of the art.
Kan Liu 0001, Bingpeng Ma, Wei Zhang 0021, Rui Huang 0001
ICCV4
2015 VoD: A novel image representation for head yaw estimation
Bingpeng Ma, Rui Huang 0001
Neurocomputing2
2015 Accurate and robust facial expressions recognition by fusing multiple sparse representation based classifiers
Yan Ouyang, Nong Sang, Rui Huang 0001
Neurocomputing3
2014 Exemplar-based linear discriminant analysis for robust object tracking
abstract
Tracking-by-detection has become an attractive tracking technique, which treats tracking as a category detection problem. However, the task in tracking is to search for a specific object, rather than an object category as in detection. In this paper, we propose a novel tracking framework based on exemplar detector rather than category detector. The proposed tracker is an ensemble of exemplar-based linear discriminant analysis (ELDA) detectors. Each detector is quite specific and discriminative, because it is trained by a single object instance and massive negatives. To improve its adaptivity, we update both object and background models. Experimental results on several challenging video sequences demonstrate the effectiveness and robustness of our tracking algorithm.
Changxin Gao, Jin-Gang Yu, Rui Huang 0001, Nong Sang
ICIP4
2014 An expressive deep model for human action parsing from a single image
abstract
This paper aims at one newly raising task in vision and multimedia research: recognizing human actions from still images. Its main challenges lie in the large variations in human poses and appearances, as well as the lack of temporal motion information. Addressing these problems, we propose to develop an expressive deep model to naturally integrate human layout and surrounding contexts for higher level action understanding from still images. In particular, a Deep Belief Net is trained to fuse information from different noisy sources such as body part detection and object detection. To bridge the semantic gap, we used manually labeled data to greatly improve the effectiveness and efficiency of the pre-training and fine-tuning stages of the DBN training. The resulting framework is shown to be robust to sometimes unreliable inputs (e.g., imprecise detections of human parts and objects), and outperforms the state-of-the-art approaches.
Zhujin Liang, Xiaolong Wang 0004, Rui Huang 0001, Liang Lin 0004
ICME3
2014 Person Search in a Scene by Jointly Modeling People Commonness and Person Uniqueness
abstract
This paper presents a novel framework for a multimedia search task: searching a person in a scene using human body appearance. Existing works mostly focus on two independent problems related to this task, i.e., people detection and person re-identification. However, a sequential combination of these two components does not solve the person search problem seamlessly for two reasons: 1) the errors in people detection are carried into person re-identification unavoidably; 2) the setting of person re-identification is different from that of person search which is essentially a verification problem. To bridge this gap, we propose a unified framework which jointly models the commonness of people (for detection) and the uniqueness of a person (for identification). We demonstrate superior performance of our approach on public benchmarks compared with the sequential combination of the state-of-the-art detection and identification algorithms.
Yuanlu Xu, Bingpeng Ma, Rui Huang 0001, Liang Lin 0004
ACM Multimedia3
2014 Image segmentation using spectral clustering of Gaussian mixture models
Shan Zeng, Rui Huang 0001, Zhen Kang, Nong Sang
Neurocomputing2
2013 Using clustering analysis to improve semi-supervised classification
Haitao Gan, Nong Sang, Rui Huang 0001, Xiaojun Tong, Zhiping Dan
Neurocomputing3
2013 A study on semi-supervised FCM algorithm
Shan Zeng, Xiaojun Tong, Nong Sang, Rui Huang 0001
Knowl. Inf. Syst.4
2012 Online Transfer Boosting for object tracking
Changxin Gao, Nong Sang, Rui Huang 0001
ICPR3
2011 Local Binary Pattern histogram based Texton learning for texture classification
abstract
Local Binary Pattern (LBP) and Texton are both widely used texture analysis techniques. In this paper we propose a patch-based texture classification method that takes advantage of both LBP and Texton. Unlike the traditional LBP methods that describe a texture with the occurrence of local binary patterns in the entire image, we compute the LBP histogram in a small region around each pixel to capture the local structure information. The texton learning method is then per- formed on these LBP histograms, resulting in a texture classification algorithm that outperforms the traditional LBP-based methods due to its preservation of local structure information. It also outperforms the traditional filtering-based texton methods due to its robustness to orientation and illumination. Experimental results on two benchmark databases validate the advantages of the proposed method.
Yonggang He, Nong Sang, Rui Huang 0001
ICIP3
2011 A Belief Propagation algorithm for bias field estimation and image segmentation
abstract
Intensity-based image segmentation is often plagued by the spatial intensity inhomogeneities (or non-uniformities) that are caused by the imperfection of the imaging devices and the varying operating conditions, also known as the bias field. We present a graphical model representation of the joint segmentation and bias field estimation problem and propose an iterative solver based on the Belief Propagation (BP) algorithm. The intractable joint inference problem of the original graphical model is decoupled into two MRF-MAP estimation problems and solved by a discrete-valued BP and a Gaussian BP, respectively and iteratively. We validate our method using both simulated and real data and show its connection to some of the classical filtering-based approaches.
Rui Huang 0001, Nong Sang, Vladimir Pavlovic 0001, Dimitris N. Metaxas
ICIP1
2011 Image segmentation via coherent clustering in L*a*b* color space
Rui Huang 0001, Nong Sang, Dapeng Luo, Qiling Tang
Pattern Recognit. Lett.1
2011 A Level Set Method for Image Segmentation in the Presence of Intensity Inhomogeneities With Application to MRI
abstract
Intensity inhomogeneity often occurs in real-world images, which presents a considerable challenge in image segmentation. The most widely used image segmentation algorithms are region-based and typically rely on the homogeneity of the image intensities in the regions of interest, which often fail to provide accurate segmentation results due to the intensity inhomogeneity. This paper proposes a novel region-based method for image segmentation, which is able to deal with intensity inhomogeneities in the segmentation. First, based on the model of images with intensity inhomogeneities, we derive a local intensity clustering property of the image intensities, and define a local clustering criterion function for the image intensities in a neighborhood of each point. This local clustering criterion function is then integrated with respect to the neighborhood center to give a global criterion of image segmentation. In a level set formulation, this criterion defines an energy in terms of the level set functions that represent a partition of the image domain and a bias field that accounts for the intensity inhomogeneity of the image. Therefore, by minimizing this energy, our method is able to simultaneously segment the image and estimate the bias field, and the estimated bias field can be used for intensity inhomogeneity correction (or bias correction). Our method has been validated on synthetic images and real images of various modalities, with desirable performance in the presence of intensity inhomogeneities. Experiments show that our method is more robust to initialization, faster and more accurate than the well-known piecewise smooth model. As an application, our method has been used for segmentation and bias correction of magnetic resonance (MR) images with promising results.
Chunming Li, Rui Huang 0001, Zhaohua Ding, Chris Gatenby, Dimitris N. Metaxas, John C. Gore
IEEE Trans. Image Process.2
2010 Saliency Based on Multi-scale Ratio of Dissimilarity
abstract
Recently, many vision applications tend to utilize saliency maps derived from input images to guide them to focus on processing salient regions in images. In this paper, we propose a simple and effective method to quantify the saliency for each pixel in images. Specially, we define the saliency for a pixel in a ratio form, where the numerator measures the number of dissimilar pixels in its center-surround and the denominator measures the total number of pixels in its center-surround. The final saliency is obtained by combining these ratios of dissimilarity over multiple scales. For images, the saliency map generated by our method not only has a high quality in resolution also looks more reasonable. Finally, we apply our saliency map to extract the salient regions in images, and compare the performance with some state-of-the-art methods over an established ground-truth which contains 1000 images.
Rui Huang 0001, Nong Sang, Leyuan Liu 0001, Qiling Tang
ICPR1
2009 Segmentation via Incremental Transductive Learning
abstract
In this paper, we propose a novel unsupervised clustering method for feature space analysis. We combine mean shift with a transductive learning method, semi-supervised discriminant analysis (SDA), in an incremental learning scheme. We use mean shift clustering to generate the class label, and use SDA to do subspace selection. Both these steps are performed alternately. Our clustering result could maintain good spatial consistency for all data in feature space. On image segmentation, we directly apply our clustering method to the L*a*b* color feature space generated from superpixels, and set each pixel with the clustering label of its superpixel. We test our image segmentation method on Berkeley image data set.
Rui Huang 0001, Nong Sang, Qiling Tang
ICIG1
2008 Approximation of salient contours in cluttered scenes
abstract
This paper proposes a new approach to describe the salient contours in cluttered scenes. No need to do the preprocessing, such as edge detection, we directly use a set of random straight line segments, as the intermediate level vision tokens, to approximate the salient contours. This line set is modeled by a stochastic framework, marked point process, in which the point denotes the center of lines, and the marker denotes the orientation and length of lines. Generic Gastalt factors of proximity and collinear continuity are embedded to constraint the geometrical inter-relations between lines. Different data likelihoods are used on synthetic and real images. Optimization is done by simulated annealing using Reversible Jump Markov chain Monte Carlo. Our results not only have a good approximation to the salient contours, also make other post-processing application more robust.
Rui Huang 0001, Nong Sang, Qiling Tang
ICPR1
2008 A Variational Level Set Approach to Segmentation and Bias Correction of Images with Intensity Inhomogeneity
Chunming Li, Rui Huang 0001, Zhaohua Ding, Chris Gatenby, Dimitris N. Metaxas, John C. Gore
MICCAI (2)2
2007 Embedded Profile Hidden Markov Models for Shape Analysis
abstract
An ideal shape model should be both invariant to global transformations and robust to local distortions. In this paper we present a new shape modeling framework that achieves both efficiently. A shape instance is described by a curvature-based shape descriptor. A Profile Hidden Markov Model (PHMM) is then built on such descriptors to represent a class of similar shapes. PHMMs are a particular type of Hidden Markov Models (HMMs) with special states and architecture that can tolerate considerable shape contour perturbations, including rigid and non-rigid deformations, occlusions, and missing parts. The sparseness of the PHMM structure provides efficient inference and learning algorithms for shape modeling and analysis. To capture the global characteristics of a class of shapes, the PHMM parameters are further embedded into a subspace that models long term spatial dependencies. The new framework can be applied to a wide range of problems, such as shape matching/registration, classification/recognition, etc. Our experimental results demonstrate the effectiveness and robustness of this new model in these different settings.
Rui Huang 0001, Vladimir Pavlovic 0001, Dimitris N. Metaxas
ICCV1
2006 Hybrid Deformable Models for Medical Segmentation and Registration
abstract
Deformable models have had great successes over the past 20 years in medical applications. We have recently developed new classes of deformable models which we term hybrid deformable models to automate the model initialization process and make improvements in segmentation and registration. In this paper we present several hybrid deformable methods we have been developing for segmentation and registration. These methods include metamorphs, a novel shape and texture integration deformable model framework and the integration of deformable models with graphical models and learning methods. We first present a framework for the robust segmentation and tracking of the heart from tagged MRI images and second applications involving brain tumor segmentation as well as brain and cardiac shape registration
Dimitris N. Metaxas, Sharon X. Huang, Rui Huang 0001, Ting Chen 0001, Leon Axel
ICARCV4
2006 A Profile Hidden Markov Model Framework for Modeling and Analysis of Shape
abstract
In this paper we propose a new framework for modeling 2D shapes. A shape is first described by a sequence of local features (e.g., curvature) of the shape boundary. The resulting description is then used to build a profile hidden Markov model (PHMM) representation of the shape. PHMMs are a particular type of hidden Markov models (HMMs) with special states and architecture that can tolerate considerable shape contour perturbations, including rigid and non-rigid deformations, occlusions and missing contour parts. Different from traditional HMM-based shape models, the sparseness of the PHMM structure allows efficient inference and learning algorithms for shape modeling and analysis. The new framework can be applied to a wide range of problems, from shape matching and classification to shape segmentation. Our experimental results show the effectiveness and robustness of this new approach in the three application domains.
Rui Huang 0001, Vladimir Pavlovic 0001, Dimitris N. Metaxas
ICIP1
2004 A Graphical Model Framework for Coupling MRFs and Deformable Models
Rui Huang 0001, Vladimir Pavlovic 0001, Dimitris N. Metaxas
CVPR (2)1
2003 Kernel-Based Nonlinear Discriminant Analysis for Face Recognition
Qingshan Liu 0001, Rui Huang 0001, Hanqing Lu, Songde Ma
J. Comput. Sci. Technol.2