EDBT 2026 Demo / reviewers in the wild / expert
Chang Huang
dblp:17/2241
· DBLP profile ↗
73ranked-venue papers
14as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 8 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 46 · 7 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 5 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DRM-Net: Explicit Residual Modelling with Subaquatic Multi-Scale Context Fusion for Underwater Image EnhancementabstractClear and high-quality underwater images are essential for marine applications, including autonomous navigation, ecological monitoring, and infrastructure inspection. However, underwater images typically suffer from severe colour distortion, low contrast, and diminished structural visibility due to wavelength-dependent attenuation, scattering, and uneven illumination conditions. Recent deep learning-based underwater image enhancement (UIE) methods primarily adopt end-to-end frameworks, directly regressing enhanced images from degraded inputs. While these approaches have achieved significant progress, they often lack explicit modeling of the degradation process, leading to limited interpretability and suboptimal recovery of fine-grained details. To address these limitations, we propose DRM-Net, an explicit residual learning framework for UIE. Rather than estimating the enhanced image directly, DRM-Net first predicts a pixel-wise Degradation Residual Map (DRM) in the perceptually uniform CIELab colour space. This map explicitly quantifies local colour, contrast, and structural degradations, thereby enabling the network to precisely reconstruct missing visual information. Furthermore, we design a lightweight Subaquatic Multi-Scale Context Fusion module, which utilizes parallel atrous convolutions with softmax-weighted feature aggregation, significantly enhancing robustness against spatially heterogeneous scattering. Trained jointly with pixel-wise DRM and VGG-based perceptual losses, DRM-Net achieves superior colour fidelity, perceptual realism, and structural detail recovery. Comprehensive experiments conducted on multiple benchmarks demonstrate that our proposed approach attains competitive quantitative results and superior qualitative visual performance compared to state-of-the-art UIE methods, while maintaining low computational overhead, making it particularly suitable for resource-constrained underwater robotic systems. Chang Huang, Zhexin Zhou, Jun Ma 0008, Jiatong Shen, Peixuan Xiong, Huayong Yang, Kaishun Wu |
AAAI | 1 |
| 2026 | mmJEPA-ECG: Cross-Posture Robust Contactless Electrocardiogram Monitoring via Millimeter Wave Radar SensingabstractContinuous cardiac monitoring during sleep is vital for detecting silent arrhythmia and other nocturnal cardiac events. While electrocardiogram (ECG) is the clinical gold standard, its reliance on electrodes and physical contact makes it intrusive for daily long-term use. Millimeter-wave (mmWave) radar offers a compelling non-contact alternative by capturing cardiac-induced chest-wall micro-vibrations. Existing radar-to-ECG methods often rely on direct waveform regression, assuming posture-stable mappings that break under natural sleep movements and obscure true cardiac rhythms. Inspired by the modality-invariant perception observed in speech and vision, we introduce mmJEPA-ECG, a physiology-guided framework for reconstructing clinical ECGs by anchoring radar sensing to invariant cardiac dynamics. It addresses two fundamental challenges: (i) disentangling robust cardiac representations from posture-induced artifacts, and (ii) generalizing ECG reconstruction across individuals under signal ambiguity. To address these challenges, Physiology-Oriented Self-Supervised Pretraining builds on a Joint Embedding Predictive Architecture (JEPA) with domain-informed masking and heart rate consistency to extract posture-robust cardiac embeddings. Conditional Diffusion-based ECG Reconstruction then generates personalized ECG waveforms through a hierarchical conditional diffusion process by spectral fidelity and denoising constraints. Extensive experiments on both public and self-collected multi-subject datasets demonstrate that our method outperforms state-of-the-art across waveform and rhythm metrics, halving R-R peak errors even under posture shifts and arrhythmic conditions. Chang Huang, Shuxin Zhong, Kaishun Wu |
AAAI | 4 |
| 2026 | BladderSense: A Wearable Ultrasound System for Continuous Bladder Monitoring in Real-World UseabstractContinuous and precise bladder volume monitoring is essential for patients with lower urinary tract dysfunction (LUTD) to support timely voiding and effective rehabilitation management. However, existing devices lack skin conformity and cannot support reliable "stick-once, long-term daily use" in real-world scenarios. We present BladderSense, a skin-conforming wireless wearable system featuring an X-shaped flexible phased-array ultrasound probe. Its development faces three key challenges: preserving beam focusing under skin deformation, maintaining accurate volume estimation despite bladder position shifts, and enabling low-power wireless transmission despite large raw data volume. To address these obstacles, we employ a suitable-frequency deep-focus design to stabilize beam quality, introduce a dual-orthogonal array with a shared geometric anchor and develop a coordinate-encoded deep learning (DL) model, together enabling bladder tracking and end-to-end volume estimation. An envelope-extraction-based compression scheme further enables Bluetooth Low Energy (BLE) transmission, supporting continuous monitoring with intermittent (1-minute) sensing. Experiments with 10 participants show that BladderSense provides accurate, robust bladder volume estimation across bladder changes, posture transitions, and dynamic daily activities, realizing dependable "stick-once, long-term monitoring" for LUTD patients. Kaixin Chen 0002, Usman Saleh Toro, Jinyu Lin, Chang Huang, Junfan Xiang, Lu Wang 0002, Huachen Cui, Lei Zhu 0003, Kaishun Wu |
MobiSys | 4 |
| 2026 | GP-BO-Driven Ensemble Learning for High-Resolution Surface Soil Moisture RetrievalabstractSurface soil moisture (SSM) plays a crucial role in hydrological processes, ecosystem dynamics, and agricultural management. Currently, high spatial resolution SSM estimation primarily relies on machine learning methods. However, in heterogeneous environments, the challenges associated with hyperparameter optimization, computational efficiency and uncertainty control compromise the robustness of these methods. To address this issue, this study introduces and evaluates an integrated strategy that combines Gaussian Process Bayesian Optimization (GP-BO) with machine learning for high-resolution SSM retrieval. The results demonstrate that the combined method significantly outperforms conventional optimization methods evaluated by the test sets from Heihe River Basin, Naqu, and Shandian River basins. Notably, the integration of GP-BO with XGBoost turned out to be the optimal combination, improving R² by 0.01–0.19 and reducing ubRMSE by 0.02–1.71 percentage points relative to conventional optimizers, while requiring the least training time for ensemble models. Furthermore, vegetation-specific GP-BO-tuned XGBoost models achieve varying degrees of accuracy improvement and reduced uncertainty across various vegetation, particularly in barren and grassland regions. These findings highlight the effectiveness of GP-BO hyperparameter optimization algorithm in reducing the uncertainties of SSM estimation in heterogeneous environments. Zuo Wang 0005, Chang Huang, Lisheng Song, Yuanhong You, Shuoqi Zhang, Zhijie Dong |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | ViG: Linear-complexity Visual Sequence Learning with Gated Linear AttentionabstractRecently, linear complexity sequence modeling networks have achieved modeling capabilities similar to Vision Transformers on a variety of computer vision tasks, while using fewer FLOPs and less memory. However, their advantage in terms of actual runtime speed is not significant. To address this issue, we introduce Gated Linear Attention (GLA) for vision, leveraging its superior hardware-awareness and efficiency. We propose direction-wise gating to capture 1D global context through bidirectional modeling and a 2D gating locality injection to adaptively inject 2D local details into 1D global context. Our hardware-aware implementation further merges forward and backward scanning into a single kernel, enhancing parallelism and reducing memory cost and latency. The proposed model, ViG, offers a favorable trade-off in accuracy, parameters, and FLOPs on ImageNet and downstream tasks, outperforming popular Transformer and CNN-based models. Bencheng Liao, Xinggang Wang, Lianghui Zhu, Chang Huang |
AAAI | 5 |
| 2025 | Omnidirectional Multi-Object TrackingabstractPanoramic imagery, with its 360° field of view, offers comprehensive information to support Multi-Object Tracking (MOT) in capturing spatial and temporal relationships of surrounding objects. However, most MOT algorithms are tailored for pinhole images with limited views, impairing their effectiveness in panoramic settings. Additionally, panoramic image distortions, such as resolution loss, geometric deformation, and uneven lighting, hinder direct adaptation of existing MOT methods, leading to significant performance degradation. To address these challenges, we propose OmniTrack, an omnidirectional MOT framework that incorporates Tracklet Management to introduce temporal cues, FlexiTrack Instances for object localization and association, and the CircularStatE Module to alleviate image and geometric distortions. This integration enables tracking in panoramic field-of-view scenarios, even under rapid sensor motion. To mitigate the lack of panoramic MOT datasets, we introduce the QuadTrack dataset—a comprehensive panoramic dataset collected by a quadruped robot, featuring diverse challenges such as panoramic fields of view, intense motion, and complex environments. Extensive experiments on the public JRDB dataset and the newly introduced QuadTrack benchmark demonstrate the state-of-the-art performance of the proposed framework. OmniTrack achieves a HOTA score of 26.92% on JRDB, representing an improvement of 3.43%, and further achieves 23.45% on QuadTrack, surpassing the baseline by 6.81%. The established dataset and source code are available at https://github.com/xifen523/OmniTrack. Hao Shi 0004, Mengfei Duan, Chang Huang, Kaiwei Wang, Kailun Yang 0001 |
CVPR | 6 |
| 2025 | DACA-Net: A Degradation-Aware Conditional Diffusion Network for Underwater Image EnhancementabstractUnderwater images typically suffer from severe colour distortions, low visibility, and reduced structural clarity due to complex optical effects such as scattering and absorption, which greatly degrade their visual quality and limit the performance of downstream visual perception tasks. Existing enhancement methods often struggle to adaptively handle diverse degradation conditions and fail to leverage underwater-specific physical priors effectively. In this paper, we propose a degradation-aware conditional diffusion model to enhance underwater images adaptively and robustly. Given a degraded underwater image as input, we first predict its degradation level using a lightweight dual-stream convolutional network, generating a continuous degradation score as semantic guidance. Based on this score, we introduce a novel conditional diffusion-based restoration network with a Swin UNet backbone, enabling adaptive noise scheduling and hierarchical feature refinement. To incorporate underwater-specific physical priors, we further propose a degradation-guided adaptive feature fusion module and a hybrid loss function that combines perceptual consistency, histogram matching, and feature-level contrast. Comprehensive experiments on benchmark datasets demonstrate that our method effectively restores underwater images with superior colour fidelity, perceptual quality, and structural details. Compared with SOTA approaches, our framework achieves significant improvements in both quantitative metrics and qualitative visual assessments. Chang Huang, Jiahang Cao, Jun Ma 0008, Kieren Yu, Cong Li 0005, Huayong Yang, Kaishun Wu |
ACM Multimedia | 1 |
| 2025 | MapTRv2: An End-to-End Framework for Online Vectorized HD Map Construction
Bencheng Liao, Shaoyu Chen, Yunchi Zhang, Bo Jiang 0011, Qian Zhang 0009, Wenyu Liu 0001, Chang Huang, Xinggang Wang |
Int. J. Comput. Vis. | 7 |
| 2025 | PolarDETR: Polar Parametrization for vision-based surround-view 3D detection
Shaoyu Chen, Xinggang Wang, Tianheng Cheng, Qian Zhang 0009, Chang Huang, Wenyu Liu 0001 |
Image Vis. Comput. | 5 |
| 2025 | An XGBoost Error Correction Model for Improving Monthly Lake Water Level Estimation on Qiangtang PlateauabstractAccurate and consistent monitoring of lake water levels is essential for understanding hydrological dynamics and climate-driven variability in remote and data-scarce regions. Satellite altimetry provides high-precision lake levels observations, but its limited spatial and temporal coverage constrain large-scale monitoring. Combining Digital Elevation Model (DEM) and remote sensing imagery offers an alternative, but the accuracy of resultant water levels is affected by the uncertainties of inherent elevation and image processing. This study proposes an XGBoost model for correcting systematic errors in inconsistent DEM-derived water levels. Twelve error-influencing parameters were incorporated, spanning lake boundary uncertainty, terrain accuracy, and area-elevation fitting errors. The results achieve a high accuracy (RMSE = 2.99 m, R2= 0.99, MAE = 0.84 m), and demonstrate robust correction performance across lakes of different sizes. We reconstructed monthly water levels time-series (2000-2021) for 965 lakes on the Qiangtang Plateau (QP) using this method. The dataset unravels divergent change trends in lake water levels on QP: a significant rise (0.12 m/a) in lakes monitored by altimetry and a slight decline (-0.028 m/a) in lakes without altimetry coverage. This highlights a potential bias when relying solely on altimetry data, as the omission of shrinking, unmonitored lakes could lead to an overestimation of regional lake expansion. We found the QP lakes are clustered into three intra-annual variation patterns, reflecting distinct hydrological and climatic influences. Uncertainties in water body extraction during frozen periods significantly reduce the accuracy of lake water levels estimations. This study provides the first plateau-scale assessment of monthly lake water levels dynamics and may offer a foundation for future research on regional hydrological changes. Yunmei Li, Jiawen Chang, Ziran Wei, Yun Chen 0010, Shiqiang Zhang, Ninglian Wang, Chang Huang |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2024 | Lane Graph as Path: Continuity-Preserving Path-Wise Modeling for Online Lane Graph Construction
Bencheng Liao, Shaoyu Chen, Bo Jiang 0011, Tianheng Cheng, Qian Zhang 0009, Wenyu Liu 0001, Chang Huang, Xinggang Wang |
ECCV (44) | 7 |
| 2024 | Focus On What Matters: Separated Models For Visual-Based RL GeneralizationabstractA primary challenge for visual-based Reinforcement Learning (RL) is to generalize effectively across unseen environments. Although previous studies have explored different auxiliary tasks to enhance generalization, few adopt image reconstruction due to concerns about exacerbating overfitting to task-irrelevant features during training. Perceiving the pre-eminence of image reconstruction in representation learning, we propose SMG (\blue{S}eparated \blue{M}odels for \blue{G}eneralization), a novel approach that exploits image reconstruction for generalization. SMG introduces two model branches to extract task-relevant and task-irrelevant representations separately from visual observations via cooperatively reconstruction. Built upon this architecture, we further emphasize the importance of task-relevant features for generalization. Specifically, SMG incorporates two additional consistency losses to guide the agent's focus toward task-relevant areas across different scenarios, thereby achieving free from overfitting. Extensive experiments in DMC demonstrate the SOTA performance of SMG in generalization, particularly excelling in video-background settings. Evaluations on robotic manipulation tasks further confirm the robustness of SMG in real-world applications. Source code is available at \url{https://anonymous.4open.science/r/SMG/}. Bowen Lv, Junqiao Zhao, Chang Huang, Hongtu Zhou, Chen Ye 0002 |
NeurIPS | 7 |
| 2024 | A Nighttime Light Based Urban Sprawl Model Revealing Reduced Electricity Intensity With Increasing Urban SizeabstractAs urbanization and industrialization continue to expand globally,urban sprawl(US) has emerged as a significant challenge to sustainable development. While research on the relationship between urban sprawl and ecological environments is well-established,the effect of urban sprawl on electricity intensity(EUS) at meso and macro scales has received limited attention due to the lack of reliable data and methodologies. To address the gap, this letter examined EUS in 204 cities in China. We began by developing an urban sprawl index using nighttime light remote sensing data for quantifying US. Then, we employed a benchmark econometric model to quantify the relationships between US and electricity intensity. Our results demonstrated that the effectiveness of using nighttime light remote sensing data as proxies for identifying urban sprawl. Moreover, we found that the US coefficient (0.181) is significantly positive, suggesting that a higher degree of urban sprawl leads to lower efficiency in electricity utilization. Heterogeneity analysis also shows that the US coefficient is the largest in small cities (5.163), followed by large cities (4.344), medium-sized cities (4.188), and megacities (0.311), demonstrating EUS basically decreases with an increase in city size. These findings offered valuable insights for Chinese policymakers in developing effective strategies for sustainable urban development and energy conservation. Kaifang Shi, Yueyan Pan, Linlin Jiang, Junru Wang, Yuanzheng Cui, Jinji Ma, Chang Huang |
IEEE Geosci. Remote. Sens. Lett. | 8 |
| 2024 | Efficient Task-Specific Feature Re-Fusion for More Accurate Object Detection and Instance SegmentationabstractFeature pyramid representations have been widely adopted in the object detection literature for better handling of variations in scale, which provide abundant information from various spatial levels for classification and localization sub-tasks. We find that inter sub-task feature disentanglement and intra sub-task feature re-fusion are crucial for final prediction performance, but are hard to be achieved simultaneously considering the computational efficiency. We find this issue can be addressed by delicate module design. In this paper, we propose an Efficient Task-specific Feature Re-fusion (ETFR) module to mitigate the dilemma. ETFR disentangles inter sub-task features, reduces the output channels of multi-scale features based on their importance and re-fuses intra sub-task features via concatenation operation. As a plug-and-play module, ETFR can remarkably and consistently improve the well-established and highly-optimized object detection and instance segmentation methods, such as RetinaNet, FCOS, BlendMask and CondInst, with neglectable extra computation cost. Extensive experiments demonstrate that ETFR has good generalization ability on various changeling datasets, including COCO, LVIS and Cityscapes. Cheng Wang 0048, Jiemin Fang, Peng Guo 0001, Rui Wu 0018, Xinggang Wang, Chang Huang, Wenyu Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2024 | Mobile anchor node assisted node collaborative localization based on light reflection in WSN
Chang Huang |
Wirel. Networks | 2 |
| 2023 | VAD: Vectorized Scene Representation for Efficient Autonomous DrivingabstractAutonomous driving requires a comprehensive understanding of the surrounding environment for reliable trajectory planning. Previous works rely on dense rasterized scene representation (e.g., agent occupancy and semantic map) to perform planning, which is computationally intensive and misses the instance-level structure information. In this paper, we propose VAD, an end-to-end vectorized paradigm for autonomous driving, which models the driving scene as a fully vectorized representation. The proposed vectorized paradigm has two significant advantages. On one hand, VAD exploits the vectorized agent motion and map elements as explicit instance-level planning constraints which effectively improves planning safety. On the other hand, VAD runs much faster than previous end-to-end planning methods by getting rid of computation-intensive rasterized representation and hand-designed post-processing steps. VAD achieves state-of-the-art end-to-end planning performance on the nuScenes dataset, outperforming the previous best method by a large margin. Our base model, VAD-Base, greatly reduces the average collision rate by 29.0% and runs 2.5× faster. Besides, a lightweight variant, VAD-Tiny, greatly improves the inference speed (up to 9.3×) while achieving comparable planning performance. We believe the excellent performance and the high efficiency of VAD are critical for the real-world deployment of an autonomous driving system. Code and models are available at https://github.com/hustvl/VAD for facilitating future research. Bo Jiang 0011, Shaoyu Chen, Bencheng Liao, Helong Zhou, Qian Zhang 0009, Wenyu Liu 0001, Chang Huang, Xinggang Wang |
ICCV | 9 |
| 2023 | MapTR: Structured Modeling and Learning for Online Vectorized HD Map Construction
Bencheng Liao, Shaoyu Chen, Xinggang Wang, Tianheng Cheng, Qian Zhang 0009, Wenyu Liu 0001, Chang Huang |
ICLR | 7 |
| 2023 | Multi-agent Decision-making at Unsignalized Intersections with Reinforcement Learning from DemonstrationsabstractIntersections are key nodes and also bottlenecks of urban road networks, so improving the traffic efficiency at intersections is beneficial to improving overall traffic throughput and mitigating traffic congestion. Previous methods such as rule-based, planning-based, and single-agent reinforcement learning usually oversimplify the policies of the surrounding vehicles and thus have difficulty modeling the complex interaction behaviors between vehicles, which limits the performance of these methods to some extent. Instead, we adopt a multi-agent reinforcement learning (MARL) approach to train and coordinate the policies of all vehicles to handle unsignalized intersection scenarios. Nevertheless, due to complex interactions between multiple agents, it is challenging to efficiently explore the environment and obtain high-reward samples. We therefore propose to pre-train the policy using demonstration data consisting of expert data and interaction data to improve the initial performance of agents and improve exploration, as well as to reduce the distributional shift between the demonstration data and the environmental interaction data. We experimentally prove that using interaction data generated by the algorithm in the demonstration data improves training stability. The proposed method enables effective exploration and greatly speeds up the training process. Chang Huang, Junqiao Zhao, Hongtu Zhou, Chen Ye 0002 |
IV | 1 |
| 2023 | How to Fine-tune the Model: Unified Model Shift and Model Bias Policy OptimizationabstractDesigning and deriving effective model-based reinforcement learning (MBRL) algorithms with a performance improvement guarantee is challenging, mainly attributed to the high coupling between model learning and policy optimization. Many prior methods that rely on return discrepancy to guide model learning ignore the impacts of model shift, which can lead to performance deterioration due to excessive model updates. Other methods use performance difference bound to explicitly consider model shift. However, these methods rely on a fixed threshold to constrain model shift, resulting in a heavy dependence on the threshold and a lack of adaptability during the training process. In this paper, we theoretically derive an optimization objective that can unify model shift and model bias and then formulate a fine-tuning process. This process adaptively adjusts the model updates to get a performance improvement guarantee while avoiding model overfitting. Based on these, we develop a straightforward algorithm USB-PO (Unified model Shift and model Bias Policy Optimization). Empirical results show that USB-PO achieves state-of-the-art performance on several challenging benchmark tasks. Junqiao Zhao, Hongtu Zhou, Chang Huang, Chen Ye 0002 |
NeurIPS | 7 |
| 2023 | Circuit as Set of PointsabstractAs the size of circuit designs continues to grow rapidly, artificial intelligence technologies are being extensively used in Electronic Design Automation (EDA) to assist with circuit design.
Placement and routing are the most time-consuming parts of the physical design process, and how to quickly evaluate the placement has become a hot research topic.
Prior works either transformed circuit designs into images using hand-crafted methods and then used Convolutional Neural Networks (CNN) to extract features, which are limited by the quality of the hand-crafted methods and could not achieve end-to-end training, or treated the circuit design as a graph structure and used Graph Neural Networks (GNN) to extract features, which require time-consuming preprocessing.
In our work, we propose a novel perspective for circuit design by treating circuit components as point clouds and using Transformer-based point cloud perception methods to extract features from the circuit. This approach enables direct feature extraction from raw data without any preprocessing, allows for end-to-end training, and results in high performance.
Experimental results show that our method achieves state-of-the-art performance in congestion prediction tasks on both the CircuitNet and ISPD2015 datasets, as well as in design rule check (DRC) violation prediction tasks on the CircuitNet dataset.
Our method establishes a bridge between the relatively mature point cloud perception methods and the fast-developing EDA algorithms, enabling us to leverage more collective intelligence to solve this task. To facilitate the research of open EDA design, source codes and pre-trained models are released at https://github.com/hustvl/circuitformer. Jialv Zou, Xinggang Wang, Wenyu Liu 0001, Qian Zhang 0009, Chang Huang |
NeurIPS | 6 |
| 2022 | AziNorm: Exploiting the Radial Symmetry of Point Cloud for Azimuth-Normalized 3D PerceptionabstractStudying the inherent symmetry of data is of great importance in machine learning. Point cloud, the most important data format for 3D environmental perception, is naturally endowed with strong radial symmetry. In this work, we exploit this radial symmetry via a divide-and-conquer strategy to boost 3D perception performance and ease optimization. We propose Azimuth Normalization (AziNorm), which normalizes the point clouds along the radial direction and eliminates the variability brought by the difference of azimuth. AziNorm can be flexibly incorporated into most LiDAR-based perception methods. To validate its effectiveness and generalization ability, we apply AziNorm in both object detection and semantic segmentation. For detection, we integrate AziNorm into two representative detection methods, the one-stage SECOND detector and the state-of-the-art two-stage PV-RCNN detector. Experiments on Waymo Open Dataset demonstrate that AziNorm improves SECOND and PV-RCNN by 7.03 mAPH and 3.01 mAPH respectively. For segmentation, we integrate AziNorm into KPConv. On SemanticKitti dataset, AziNorm improves KPConv by 1.6/1.1 mIoU on val/test set. Besides, AziNorm remarkably improves data efficiency and accelerates convergence, reducing the requirement of data amounts or training epochs by an order of magnitude. SECOND w/ AziNorm can significantly outperform fully trained vanilla SECOND, even trained with only 10% data or 10% epochs. Code and models are available at https://github.com/hustvl/AziNorm. Shaoyu Chen, Xinggang Wang, Tianheng Cheng, Qian Zhang 0009, Chang Huang, Wenyu Liu 0001 |
CVPR | 6 |
| 2022 | Sparse Instance Activation for Real-Time Instance SegmentationabstractIn this paper, we propose a conceptually novel, efficient, and fully convolutional framework for real-time instance segmentation. Previously, most instance segmentation methods heavily rely on object detection and perform mask prediction based on bounding boxes or dense centers. In contrast, we propose a sparse set of instance activation maps, as a new object representation, to high-light informative regions for each foreground object. Then instance-level features are obtained by aggregating features according to the highlighted regions for recognition and segmentation. Moreover, based on bipartite matching, the instance activation maps can predict objects in a one-to-one style, thus avoiding non-maximum suppression (NMS) in post-processing. Owing to the simple yet effective designs with instance activation maps, SparseInst has extremely fast inference speed and achieves 40 FPS and 37.9 AP on the COCO benchmark, which significantly out-performs the counterparts in terms of speed and accuracy. Code and models are available at https://github.com/hustvl/SparseInst. Tianheng Cheng, Xinggang Wang, Shaoyu Chen, Qian Zhang 0009, Chang Huang, Zhaoxiang Zhang 0001, Wenyu Liu 0001 |
CVPR | 6 |
| 2022 | Fusing Landsat-8, Sentinel-1, and Sentinel-2 Data for River Water Mapping Using Multidimensional Weighted Fusion MethodabstractRiver water extent is critical for understanding river discharge or its hydrological conditions. Although numerous methods have been proposed to map river water from either optical or synthetic aperture radar (SAR) remotely sensed images, uncertainties still exist broadly. In this study, we developed an image fusion method that integrates Landsat-8, Sentinel-1 and Sentinel-2 images simultaneously for river water mapping with two major steps. Firstly, a posterior probability support vector machine model was adopted to generate water probability maps from each individual image; and second, a Multi-dimensional Weighted Fusion Method (MDWFM) was developed to fuse these probability maps. Four reaches with different characteristics were selected as case study sites. High resolution aerial images were acquired and used as the reference to evaluate our results. We found the fusion process not only improves the quality of river water mapping, but also excludes the cloud interference. The fused river water maps become more reliable after the conflicts from difference images being solved by the proposed MDWFM method that contains a proportional conflict redistribution rule. The weighted root mean square difference was reduced to 0.066, and the Area Under the ROC curve reached up to 0.984. The Critical Success Index, Kappa Coefficient, and F-measure reached up to 0.810, 0.836 and 0.895, respectively. These stable and accurate river extent mapping results obtained through fusing multiple images with high spatial resolution (10 m) and short revisit interval (0.4~4.4 days) are of great significance for enriching the data and methodology of hydrological studies. Qihang Liu, Shiqiang Zhang, Ninglian Wang, Yisen Ming, Chang Huang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | Image Registration Improved by Generative Adversarial Networks
Shiyan Jiang, Ci Wang, Chang Huang |
MMM (2) | 3 |
| 2021 | EAT-NAS: elastic architecture transfer for accelerating large-scale neural architecture search
Jiemin Fang, Yukang Chen, Xinbang Zhang, Qian Zhang 0009, Chang Huang, Gaofeng Meng, Wenyu Liu 0001, Xinggang Wang |
Sci. China Inf. Sci. | 5 |
| 2021 | Real-time and accurate object detection in compressed video by long short-term feature aggregation
Xinggang Wang, Zhaojin Huang, Bencheng Liao, Lichao Huang, Yongchao Gong, Chang Huang |
Comput. Vis. Image Underst. | 6 |
| 2020 | Diversity Transfer Network for Few-Shot LearningabstractFew-shot learning is a challenging task that aims at training a classifier for unseen classes with only a few training examples. The main difficulty of few-shot learning lies in the lack of intra-class diversity within insufficient training samples. To alleviate this problem, we propose a novel generative framework, Diversity Transfer Network (DTN), that learns to transfer latent diversities from known categories and composite them with support features to generate diverse samples for novel categories in feature space. The learning problem of the sample generation (i.e., diversity transfer) is solved via minimizing an effective meta-classification loss in a single-stage network, instead of the generative loss in previous works. Besides, an organized auxiliary task co-training over known categories is proposed to stabilize the meta-training process of DTN. We perform extensive experiments and ablation studies on three datasets, i.e., miniImageNet, CIFAR100 and CUB. The results show that DTN, with single-stage training and faster convergence speed, obtains the state-of-the-art results among the feature generation based few-shot learning methods. Code and supplementary material are available at: https://github.com/Yuxin-CV/DTN. Xinggang Wang, Yifeng Geng, Chang Huang, Wenyu Liu 0001, Bo Wang 0044 |
AAAI | 7 |
| 2020 | RDSNet: A New Deep Architecture forReciprocal Object Detection and Instance SegmentationabstractObject detection and instance segmentation are two fundamental computer vision tasks. They are closely correlated but their relations have not yet been fully explored in most previous work. This paper presents RDSNet, a novel deep architecture for reciprocal object detection and instance segmentation. To reciprocate these two tasks, we design a two-stream structure to learn features on both the object level (i.e., bounding boxes) and the pixel level (i.e., instance masks) jointly. Within this structure, information from the two streams is fused alternately, namely information on the object level introduces the awareness of instance and translation variance to the pixel level, and information on the pixel level refines the localization accuracy of objects on the object level in return. Specifically, a correlation module and a cropping module are proposed to yield instance masks, as well as a mask based boundary refinement module for more accurate bounding boxes. Extensive experimental analyses and comparisons on the COCO dataset demonstrate the effectiveness and efficiency of RDSNet. The source code is available at https://github.com/wangsr126/RDSNet. Shaoru Wang, Yongchao Gong, Junliang Xing, Lichao Huang, Chang Huang, Weiming Hu 0004 |
AAAI | 5 |
| 2020 | Unsupervised domain adaptive re-identification: Theory and practice
Liangchen Song, Cheng Wang 0048, Lefei Zhang, Bo Du 0001, Qian Zhang 0009, Chang Huang, Xinggang Wang |
Pattern Recognit. | 6 |
| 2019 | RENAS: Reinforced Evolutionary Neural Architecture SearchabstractNeural Architecture Search (NAS) is an important yet challenging task in network design due to its high computational consumption. To address this issue, we propose the Reinforced Evolutionary Neural Architecture Search (RENAS), which is an evolutionary method with reinforced mutation for NAS. Our method integrates reinforced mutation into an evolution algorithm for neural architecture exploration, in which a mutation controller is introduced to learn the effects of slight modifications and make mutation actions. The reinforced mutation controller guides the model population to evolve efficiently. Furthermore, as child models can inherit parameters from their parents during evolution, our method requires very limited computational resources. In experiments, we conduct the proposed search method on CIFAR-10 and obtain a powerful network architecture, RENASNet. This architecture achieves a competitive result on CIFAR-10. The explored network architecture is transferable to ImageNet and achieves a new state-of-the-art accuracy, i.e., 75.7% top-1 accuracy with 5.36M parameters on mobile ImageNet. We further test its performance on semantic segmentation with DeepLabv3 on the PASCAL VOC. RENASNet outperforms MobileNet-v1, MobileNet-v2 and NASNet. It achieves 75.83% mIOU without being pretrained on COCO. Yukang Chen, Gaofeng Meng, Qian Zhang 0009, Shiming Xiang, Chang Huang, Lisen Mu, Xinggang Wang |
CVPR | 5 |
| 2019 | Mask Scoring R-CNNabstractLetting a deep network be aware of the quality of its own predictions is an interesting yet important problem. In the task of instance segmentation, the confidence of instance classification is used as mask quality score in most instance segmentation frameworks. However, the mask quality, quantified as the IoU between the instance mask and its ground truth, is usually not well correlated with classification score. In this paper, we study this problem and propose Mask Scoring R-CNN which contains a network block to learn the quality of the predicted instance masks. The proposed network block takes the instance feature and the corresponding predicted mask together to regress the mask IoU. The mask scoring strategy calibrates the misalignment between mask quality and mask score, and improves instance segmentation performance by prioritizing more accurate mask predictions during COCO AP evaluation. By extensive evaluations on the COCO dataset, Mask Scoring R-CNN brings consistent and noticeable gain with different models and outperforms the state-of-the-art Mask R-CNN. We hope our simple and effective approach will provide a new direction for improving instance segmentation. The source code of our method is available at \url{https://github.com/zjhuang22/maskscoring_rcnn}. Zhaojin Huang, Lichao Huang, Yongchao Gong, Chang Huang, Xinggang Wang |
CVPR | 4 |
| 2019 | CCNet: Criss-Cross Attention for Semantic SegmentationabstractFull-image dependencies provide useful contextual information to benefit visual understanding problems. In this work, we propose a Criss-Cross Network (CCNet) for obtaining such contextual information in a more effective and efficient way. Concretely, for each pixel, a novel criss-cross attention module in CCNet harvests the contextual information of all the pixels on its criss-cross path. By taking a further recurrent operation, each pixel can finally capture the full-image dependencies from all pixels. Overall, CCNet is with the following merits: 1) GPU memory friendly. Compared with the non-local block, the proposed recurrent criss-cross attention module requires 11x less GPU memory usage. 2) High computational efficiency. The recurrent criss-cross attention significantly reduces FLOPs by about 85% of the non-local block in computing full-image dependencies. 3) The state-of-the-art performance. We conduct extensive experiments on popular semantic segmentation benchmarks including Cityscapes, ADE20K, and instance segmentation benchmark COCO. In particular, our CCNet achieves the mIoU score of 81.4 and 45.22 on Cityscapes test set and ADE20K validation set, respectively, which are the new state-of-the-art results. The source code is available at https://github.com/speedinghzl/CCNet. Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, Wenyu Liu 0001 |
ICCV | 4 |
| 2019 | Mapping Spatio-Temporal Dynamics of Rainstorms in Recent 20 Years of China Using TRMM DataabstractSatellite missions such as Tropical Rainfall Measurement Mission (TRMM) and Global Precipitation Mission (GPM) have been collecting a large volume of precipitation data in recent years. These data provide large scale and long-term continuous precipitation information which is useful for climate change studies. Extreme precipitation events, or so called rainstorms, are a major driven for some disasters, such as urban flooding and mountain torrents. This study proposes an efficient tool of identifying rainstorms automatically from grid based precipitation dataset. Using this tool, rainstorm events in recent 20 years in China were extracted from time series of TRMM 3B42-V7 product. The seasonal and interannual variation of rainstorms, together with their occurrences, seasonality are mapped and analyzed. Intensity of the biggest rainstorms in different regions are compared and analyzed. Chang Huang, Shiqiang Zhang, Zucheng Wang |
IGARSS | 1 |
| 2019 | Learning Deep Decentralized Policy Network by Collective Rewards for Real-Time Combat GameabstractThe task of real-time combat game is to coordinate multiple units to defeat their enemies controlled by the given opponent in a real-time combat scenario. It is difficult to design a high-level Artificial Intelligence (AI) program for such a task due to its extremely large state-action space and real-time requirements. This paper formulates this task as a collective decentralized partially observable Markov decision process, and designs a Deep Decentralized Policy Network (DDPN) to model the polices. To train DDPN effectively, a novel two-stage learning algorithm is proposed which combines imitation learning from opponent and reinforcement learning by no-regret dynamics. Extensive experimental results on various combat scenarios indicate that proposed method can defeat different opponent models and significantly outperforms many state-of-the-art approaches. Peixi Peng, Junliang Xing, Lili Cao, Lisen Mu, Chang Huang |
IJCAI | 5 |
| 2019 | Enhanced Super-Resolution Mapping of Urban Floods Based on the Fusion of Support Vector Machine and General Regression Neural NetworkabstractSuper-resolution mapping of urban flood (SMUF) is one of the hotspots in remote sensing and urban environment research. In this letter, a new SMUF method based on the fusion of support vector machine and general regression neural network (FSVMGRNN) was proposed to achieve enhanced performance. An SVM-SMUF algorithm was developed and a fusion criterion was formulated. Then, the FSVMGRNN-SMUF algorithm was developed. The results of FSVMGRNN-SMUF were evaluated using Landsat 8 OLI imagery of two representative cities in China. FSVMGRNN-SMUF yielded the most accurate SMUF results among the five SMUF methods according to visual comparisons and quantitative comparisons. The mapping accuracy of FSVMGRNN-SMUF related to the kernel functions was also analyzed and discussed. The results of this letter will help to boost practical applications of median-low resolution remote sensing images in urban flooding mapping, and to strengthen the means for monitoring and assessing urban flooding disasters. Linyi Li 0002, Yun Chen 0010, Tingbao Xu, Kaifang Shi, Chang Huang, Binbin Lu, Lingkui Meng |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2018 | Mancs: A Multi-task Attentional Network with Curriculum Sampling for Person Re-Identification
Cheng Wang 0048, Qian Zhang 0009, Chang Huang, Wenyu Liu 0001, Xinggang Wang |
ECCV (4) | 3 |
| 2018 | SafeNet: Scale-normalization and Anchor-based Feature Extraction Network for Person Re-identificationabstractPerson Re-identification (ReID) is a challenging retrieval task that requires matching a person's image across non-overlapping camera views. The quality of fulfilling this task is largely determined on the robustness of the features that are used to describe the person. In this paper, we show the advantage of jointly utilizing multi-scale abstract information to learn powerful features over full body and parts. A scale normalization module is proposed to balance different scales through residual-based integration. To exploit the information hidden in non-rigid body parts, we propose an anchor-based method to capture the local contents by stacking convolutions of kernels with various aspect ratios, which focus on different spatial distributions. Finally, a well-defined framework is constructed for simultaneously learning the representations of both full body and parts. Extensive experiments conducted on current challenging large-scale person ReID datasets, including Market1501, CUHK03 and DukeMTMC, demonstrate that our proposed method achieves the state-of-the-art results. Kun Yuan 0003, Qian Zhang 0009, Chang Huang, Shiming Xiang, Chunhong Pan |
IJCAI | 3 |
| 2016 | CNN-RNN: A Unified Framework for Multi-label Image ClassificationabstractWhile deep convolutional neural networks (CNNs) have shown a great success in single-label image classification, it is important to note that real world images generally contain multiple labels, which could correspond to different objects, scenes, actions and attributes in an image. Traditional approaches to multi-label image classification learn independent classifiers for each category and employ ranking or thresholding on the classification results. These techniques, although working well, fail to explicitly exploit the label dependencies in an image. In this paper, we utilize recurrent neural networks (RNNs) to address this problem. Combined with CNNs, the proposed CNN-RNN framework learns a joint image-label embedding to characterize the semantic label dependency as well as the image-label relevance, and it can be trained end-to-end from scratch to integrate both information in a unified framework. Experimental results on public benchmark datasets demonstrate that the proposed architecture achieves better performance than the state-of-the-art multi-label classification models. Jiang Wang 0001, Yi Yang 0007, Junhua Mao, Zhiheng Huang, Chang Huang, Wei Xu 0017 |
CVPR | 5 |
| 2016 | Surface water change detection using change vector analysisabstractMonitoring the dynamics of surface water using remote sensing technology is an essential research topic in many research areas. Change vector analysis (CVA) is a change detection method by constructing a vector of change based on multi-temporal images. It has been widely applied in land use and land cover change detection. Considering the specialty of surface water's spectral characteristics, it is anticipated that the CVA method can be revised and improved specially for monitoring surface water change. This study utilizes three Landsat images at different time phases for testing this proposed method, and uses traditional classification results for evaluation. Evaluation results demonstrate that surface water dynamic has been detected successfully using the CVA method. The proposed method along with its calibrated criteria can be applied for detecting surface water change between any other two time phases for this area. This can therefore be helpful for quick monitoring of water dynamic. Chang Huang, Xiaoyu Zan, Shiqiang Zhang |
IGARSS | 1 |
| 2015 | Multi-objective convolutional learning for face labelingabstractThis paper formulates face labeling as a conditional random field with unary and pairwise classifiers. We develop a novel multi-objective learning method that optimizes a single unified deep convolutional network with two distinct non-structured loss functions: one encoding the unary label likelihoods and the other encoding the pairwise label dependencies. Moreover, we regularize the network by using a nonparametric prior as new input channels in addition to the RGB image, and show that significant performance improvements can be achieved with a much smaller network size. Experiments on both the LFW and Helen datasets demonstrate state-of-the-art results of the proposed algorithm, and accurate labeling results on challenging images can be obtained by the proposed algorithm for real-world applications. Sifei Liu, Jimei Yang, Chang Huang, Ming-Hsuan Yang 0001 |
CVPR | 3 |
| 2015 | Deep multiple instance learning for image classification and auto-annotationabstractThe recent development in learning deep representations has demonstrated its wide applications in traditional vision tasks like classification and detection. However, there has been little investigation on how we could build up a deep learning framework in a weakly supervised setting. In this paper, we attempt to model deep learning in a weakly supervised learning (multiple instance learning) framework. In our setting, each image follows a dual multi-instance assumption, where its object proposals and possible text annotations can be regarded as two instance sets. We thus design effective systems to exploit the MIL property with deep learning strategies from the two ends; we also try to jointly learn the relationship between object and annotation proposals. We conduct extensive experiments and prove that our weakly supervised deep learning framework not only achieves convincing performance in vision tasks including classification and image annotation, but also extracts reasonable region-keyword pairs with little supervision, on both widely used benchmarks like PASCAL VOC and MIT Indoor Scene 67, and also a dataset for image-and patch-level annotations. Jiajun Wu 0001, Yinan Yu, Chang Huang |
CVPR | 3 |
| 2015 | Learning from massive noisy labeled data for image classificationabstractLarge-scale supervised datasets are crucial to train convolutional neural networks (CNNs) for various computer vision problems. However, obtaining a massive amount of well-labeled data is usually very expensive and time consuming. In this paper, we introduce a general framework to train CNNs with only a limited number of clean labels and millions of easily obtained noisy labels. We model the relationships between images, class labels and label noises with a probabilistic graphical model and further integrate it into an end-to-end deep learning system. To demonstrate the effectiveness of our approach, we collect a large-scale real-world clothing classification dataset with both noisy and clean labels. Experiments on this dataset indicate that our approach can better correct the noisy labels and improves the performance of trained CNNs. Tong Xiao 0003, Chang Huang, Xiaogang Wang 0001 |
CVPR | 4 |
| 2015 | Conditional Random Fields as Recurrent Neural NetworksabstractPixel-level labelling tasks, such as semantic segmentation, play a central role in image understanding. Recent approaches have attempted to harness the capabilities of deep learning techniques for image recognition to tackle pixel-level labelling tasks. One central issue in this methodology is the limited capacity of deep learning techniques to delineate visual objects. To solve this problem, we introduce a new form of convolutional neural network that combines the strengths of Convolutional Neural Networks (CNNs) and Conditional Random Fields (CRFs)-based probabilistic graphical modelling. To this end, we formulate Conditional Random Fields with Gaussian pairwise potentials and mean-field approximate inference as Recurrent Neural Networks. This network, called CRF-RNN, is then plugged in as a part of a CNN to obtain a deep network that has desirable properties of both CNNs and CRFs. Importantly, our system fully integrates CRF modelling with CNNs, making it possible to train the whole deep network end-to-end with the usual back-propagation algorithm, avoiding offline post-processing methods for object delineation. We apply the proposed method to the problem of semantic image segmentation, obtaining top results on the challenging Pascal VOC 2012 segmentation benchmark. Shuai Zheng 0001, Sadeep Jayasumana, Bernardino Romera-Paredes, Vibhav Vineet, Zhizhong Su, Dalong Du, Chang Huang, Philip Torr 0001 |
ICCV | 7 |
| 2015 | Look and Think Twice: Capturing Top-Down Visual Attention with Feedback Convolutional Neural NetworksabstractWhile feedforward deep convolutional neural networks (CNNs) have been a great success in computer vision, it is important to note that the human visual cortex generally contains more feedback than feedforward connections. In this paper, we will briefly introduce the background of feedbacks in the human visual cortex, which motivates us to develop a computational feedback mechanism in deep neural networks. In addition to the feedforward inference in traditional neural networks, a feedback loop is introduced to infer the activation status of hidden layer neurons according to the "goal" of the network, e.g., high-level semantic labels. We analogize this mechanism as "Look and Think Twice." The feedback networks help better visualize and understand how deep neural networks work, and capture visual attention on expected objects, even in images with cluttered background and multiple objects. Experiments on ImageNet dataset demonstrate its effectiveness in solving tasks such as image classification and object localization. Chunshui Cao, Xianming Liu 0005, Yi Yang 0007, Yinan Yu, Jiang Wang 0001, Zilei Wang, Yongzhen Huang, Liang Wang 0001, Chang Huang, Wei Xu 0017, Deva Ramanan, Thomas S. Huang |
ICCV | 9 |
| 2015 | A Deep Visual Correspondence Embedding Model for Stereo Matching CostsabstractThis paper presents a data-driven matching cost for stereo matching. A novel deep visual correspondence embedding model is trained via Convolutional Neural Network on a large set of stereo images with ground truth disparities. This deep embedding model leverages appearance data to learn visual similarity relationships between corresponding image patches, and explicitly maps intensity values into an embedding feature space to measure pixel dissimilarities. Experimental results on KITTI and Middlebury data sets demonstrate the effectiveness of our model. First, we prove that the new measure of pixel dissimilarity outperforms traditional matching costs. Furthermore, when integrated with a global stereo framework, our method ranks top 3 among all two-frame algorithms on the KITTI benchmark. Finally, cross-validation results show that our model is able to make correct predictions for unseen data which are outside of its labeled training set. Zhuoyuan Chen, Liang Wang 0001, Yinan Yu, Chang Huang |
ICCV | 5 |
| 2015 | Text Flow: A Unified Text Detection System in Natural Scene ImagesabstractThe prevalent scene text detection approach follows four sequential steps comprising character candidate detection, false character candidate removal, text line extraction, and text line verification. However, errors occur and accumulate throughout each of these sequential steps which often lead to low detection performance. To address these issues, we propose a unified scene text detection system, namely Text Flow, by utilizing the minimum cost (min-cost) flow network model. With character candidates detected by cascade boosting, the min-cost flow network model integrates the last three sequential steps into a single process which solves the error accumulation problem at both character level and text line level effectively. The proposed technique has been tested on three public datasets, i.e, ICDAR2011 dataset, ICDAR2013 dataset and a multilingual dataset and it outperforms the state-of-the-art methods on all three datasets with much higher recall and F-score. The good performance on the multilingual dataset shows that the proposed technique can be used for the detection of texts in different languages. Shangxuan Tian, Yifeng Pan, Chang Huang, Shijian Lu, Chew Lim Tan |
ICCV | 3 |
| 2013 | A dem-based modified pixel swapping algorithm for floodplain inundation mapping at subpixel scaleabstractSubpixel mapping is a promising way to increase the spatial resolution of classification results from images that have coarse spatial resolution but high temporal resolution. Existing subpixel mapping methods are not adequate for mapping linear features such as floodplain inundation. This study modified the commonly used pixel swapping (PS) algorithm and one of its derivatives, the linearised pixel swapping (LPS) algorithm, by employing finer resolution Digital Elevation Model (DEM) data. Results of a case study show that the modified method performs better than both the PS and LPS algorithms. It improves the accuracy and the Kappa coefficient by 4.78% and 0.11 in comparison with the PS algorithm. The spatial pattern of the inundation reveals fewer breakpoints and errors along the river channels. It is hoped that the proposed method will broaden the application of coarse resolution images in flood inundation detection. Chang Huang, Yun Chen 0010 |
IGARSS | 1 |
| 2013 | Multiple Target Tracking by Learning-Based Hierarchical Association of Detection ResponsesabstractWe propose a hierarchical association approach to multiple target tracking from a single camera by progressively linking detection responses into longer track fragments (i.e., tracklets). Given frame-by-frame detection results, a conservative dual-threshold method that only links very similar detection responses between consecutive frames is adopted to generate initial tracklets with minimum identity switches. Further association of these highly fragmented tracklets at each level of the hierarchy is formulated as a Maximum A Posteriori (MAP) problem that considers initialization, termination, and transition of tracklets as well as the possibility of them being false alarms, which can be efficiently computed by the Hungarian algorithm. The tracklet affinity model, which measures the likelihood of two tracklets belonging to the same target, is a linear combination of automatically learned weak nonparametric models upon various features, which is distinct from most of previous work that relies on heuristic selection of parametric models and manual tuning of their parameters. For this purpose, we develop a novel bag ranking method and train the crucial tracklet affinity models by the boosting algorithm. This bag ranking method utilizes the soft max function to relax the oversufficient objective function used by the conventional instance ranking method. It provides a tighter upper bound of empirical errors in distinguishing correct associations from the incorrect ones, and thus yields more accurate tracklet affinity models for the tracklet association problem. We apply this approach to the challenging multiple pedestrian tracking task. Systematic experiments conducted on two real-life datasets show that the proposed approach outperforms previous state-of-the-art algorithms in terms of tracking accuracy, in particular, considerably reducing fragmentations and identity switches. Chang Huang, Yuan Li 0022, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | Beyond spatial pyramids: Receptive field learning for pooled image featuresabstractIn this paper we examine the effect of receptive field designs on classification accuracy in the commonly adopted pipeline of image classification. While existing algorithms usually use manually defined spatial regions for pooling, we show that learning more adaptive receptive fields increases performance even with a significantly smaller codebook size at the coding layer. To learn the optimal pooling parameters, we adopt the idea of over-completeness by starting with a large number of receptive field candidates, and train a classifier with structured sparsity to only use a sparse subset of all the features. An efficient algorithm based on incremental feature selection and retraining is proposed for fast learning. With this method, we achieve the best published performance on the CIFAR-10 dataset, using a much lower dimensional feature space than previous methods. Yangqing Jia, Chang Huang, Trevor Darrell |
CVPR | 2 |
| 2012 | Unsupervised incremental learning for improved object detection in a videoabstractMost common approaches for object detection collect thousands of training examples and train a detector in an offline setting, using supervised learning methods, with the objective of obtaining a generalized detector that would give good performance on various test datasets. However, when an offline trained detector is applied on challenging test datasets, it may fail in some cases by not being able to detect some objects or by producing false alarms. We propose an unsupervised multiple instance learning (MIL) based incremental solution to deal with this issue. We introduce an MIL loss function for Real Adaboost and present a tracking based effective unsupervised online sample collection mechanism to collect the online samples for incremental learning. Experiments demonstrate the effectiveness of our approach by improving the performance of a state of the art offline trained detector on the challenging datasets for pedestrian category. Pramod Sharma, Chang Huang, Ramakant Nevatia |
CVPR | 2 |
| 2012 | Efficient incremental learning of boosted classifiers for object detection
Pramod Sharma, Chang Huang, Ramakant Nevatia |
ICPR | 2 |
| 2011 | Learning affinities and dependencies for multi-target tracking using a CRF modelabstractWe propose a learning-based Conditional Random Field (CRF) model for tracking multiple targets by progressively associating detection responses into long tracks. Tracking task is transformed into a data association problem, and most previous approaches developed heuristical parametric models or learning approaches for evaluating independent affinities between track fragments (tracklets). We argue that the independent assumption is not valid in many cases, and adopt a CRF model to consider both tracklet affinities and dependencies among them, which are represented by unary term costs and pairwise term costs respectively. Unlike previous methods, we learn the best global associations instead of the best local affinities between tracklets, and transform the task of finding the best association into an energy minimization problem. A RankBoost algorithm is proposed to select effective features for estimation of term costs in the CRF model, so that better associations have lower costs. Our approach is evaluated on challenging pedestrian data sets, and are compared with state-of-art methods. Experiments show effectiveness of our algorithm as well as improvement in tracking performance. Bo Yang 0008, Chang Huang, Ramakant Nevatia |
CVPR | 2 |
| 2011 | Segmentation of objects in a detection window by Nonparametric Inhomogeneous CRFs
Bo Yang 0008, Chang Huang, Ramakant Nevatia |
Comput. Vis. Image Underst. | 2 |
| 2010 | High performance object detection by collaborative learning of Joint Ranking of Granules featuresabstractObject detection remains an important but challenging task in computer vision. We present a method that combines high accuracy with high efficiency. We adopt simplified forms of APCF features [3], which we term Joint Ranking of Granules (JRoG) features; the features consists of discrete values by uniting binary ranking results of pair-wise granules in the image. We propose a novel collaborative learning method for JRoG features, which consists of a Simulated Annealing (SA) module and an incremental feature selection module. The two complementary modules collaborate to efficiently search the formidably large JRoG feature space for discriminative features, which are fed into a boosted cascade for object detection. To cope with occlusions in crowded environments, we employ the strategy of part based detection, as in [19] but propose a new dynamic search method to improve the Bayesian combination of the part detection results. Experiments on several challenging data sets show that our approach achieves not only considerable improvement in detection accuracy but also major improvements in computational efficiency; on a Xeon 3GHz computer, with only a single thread, it can process a million scanning windows per second, sufficing for many practical real-time detection tasks. Chang Huang, Ramakant Nevatia |
CVPR | 1 |
| 2010 | Multi-target tracking by on-line learned discriminative appearance modelsabstractWe present an approach for online learning of discriminative appearance models for robust multi-target tracking in a crowded scene from a single camera. Although much progress has been made in developing methods for optimal data association, there has been comparatively less work on the appearance models, which are key elements for good performance. Many previous methods either use simple features such as color histograms, or focus on the discriminability between a target and the background which does not resolve ambiguities between the different targets. We propose an algorithm for learning a discriminative appearance model for different targets. Training samples are collected online from tracklets within a time sliding window based on some spatial-temporal constraints; this allows the models to adapt to target instances. Learning uses an Ad-aBoost algorithm that combines effective image descriptors and their corresponding similarity measurements. We term the learned models as OLDAMs. Our evaluations indicate that OLDAMs have significantly higher discrimination between different targets than conventional holistic color histograms, and when integrated into a hierarchical association framework, they help improve the tracking accuracy, particularly reducing the false alarms and identity switches. Cheng-Hao Kuo, Chang Huang, Ramakant Nevatia |
CVPR | 2 |
| 2010 | Inter-camera Association of Multi-target Tracks by On-Line Learned Appearance Affinity Models
Cheng-Hao Kuo, Chang Huang, Ramakant Nevatia |
ECCV (1) | 2 |
| 2010 | Efficient Inference with Multiple Heterogeneous Part Detectors for Human Pose Estimation
Vivek K. Singh 0002, Ramakant Nevatia, Chang Huang |
ECCV (3) | 3 |
| 2009 | Learning to associate: HybridBoosted multi-target tracker for crowded sceneabstractWe propose a learning-based hierarchical approach of multi-target tracking from a single camera by progressively associating detection responses into longer and longer track fragments (tracklets) and finally the desired target trajectories. To define tracklet affinity for association, most previous work relies on heuristically selected parametric models; while our approach is able to automatically select among various features and corresponding non-parametric models, and combine them to maximize the discriminative power on training data by virtue of a HybridBoost algorithm. A hybrid loss function is used in this algorithm because the association of tracklet is formulated as a joint problem of ranking and classification: the ranking part aims to rank correct tracklet associations higher than other alternatives; the classification part is responsible to reject wrong associations when no further association should be done. Experiments are carried out by tracking pedestrians in challenging datasets. We compare our approach with state-of-the-art algorithms to show its improvement in terms of tracking accuracy. Yuan Li 0022, Chang Huang, Ramakant Nevatia |
CVPR | 2 |
| 2009 | Extensive articulated human detection by voting Cluster Boosted TreeabstractOur goal is to detect people in highly articulated poses, including bending, crouching, etc. Such formidable diversity in human poses makes detection much more difficult than for pedestrian poses. ¿Divide-and-conquer¿ is a favorable strategy for detecting objects with large intra class variations, which splits object instances into several subcategories and trains relatively simple classifiers for each sub-category. We propose a novel sample split method, which benefits the learning results of articulated humans. We adopt the cluster boosted tree (CBT) structure to automatically decide when a split should be triggered. Unlike the simple k-means used in CBT for sample split, our approach aims at minimizing the training loss after the split. Since this minimization is an NP-hard problem, we design a heuristic algorithm, in which we find optimal sample divisions according to each single feature, and then make compromises to get a final division by a voting-like process. We name our training method as voting cluster boosted tree (VCBT). Furthermore, to avoid large background area in training samples, we first cluster samples according to their width/height ratios, and then train a VCBT for each subset. We conduct an experiment on 17 infrared surveillance video clips, report superior performance compared with previous human detection methods, and show how our approach benefits the learning results by reducing training loss. Bo Yang 0008, Chang Huang, Ramakant Nevatia |
WACV | 2 |
| 2008 | Robust Object Tracking by Hierarchical Association of Detection Responses
Chang Huang, Bo Wu 0001, Ramakant Nevatia |
ECCV (2) | 1 |
| 2007 | Incremental Learning of Boosted Face DetectorabstractIn recent years, boosting has been successfully applied to many practical problems in pattern recognition and computer vision fields such as object detection and tracking. As boosting is an offline training process with beforehand collected data, once learned, it cannot make use of any newly arriving ones. However, an offline boosted detector is to be exploited online and inevitably there must be some special cases that are not covered by those beforehand collected training data. As a result, the inadaptable detector often performs badly in diverse and changeful environments which are ordinary for many real-life applications. To alleviate this problem, this paper proposes an incremental learning algorithm to effectively adjust a boosted strong classifier with domain-partitioning weak hypotheses to online samples, which adopts a novel approach to efficient estimation of training losses received from offline samples. By this means, the offline learned general-purpose detectors can be adapted to special online situations at a low extra cost, and still retains good generalization ability for common environments. The experiments show convincing results of our incremental learning approach on challenging face detection problems with partial occlusions and extreme illuminations. Chang Huang, Haizhou Ai, Takayoshi Yamashita, Shihong Lao, Masato Kawade |
ICCV | 1 |
| 2007 | High-Performance Rotation Invariant Multiview Face DetectionabstractRotation invariant multiview face detection (MVFD) aims to detect faces with arbitrary rotation-in-plane (RIP) and rotation-off-plane (ROP) angles in still images or video sequences. MVFD is crucial as the first step in automatic face processing for general applications since face images are seldom upright and frontal unless they are taken cooperatively. In this paper, we propose a series of innovative methods to construct a high-performance rotation invariant multiview face detector, including the Width-First-Search (WFS) tree detector structure, the Vector Boosting algorithm for learning vector-output strong classifiers, the domain-partition-based weak learning method, the sparse feature in granular space, and the heuristic search for sparse feature selection. As a result of that, our multiview face detector achieves low computational complexity, broad detection scope, and high detection accuracy on both standard testing sets and real-life images. Chang Huang, Haizhou Ai, Yuan Li 0022, Shihong Lao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | VLSI Design of a High-Speed and Area-Efficient JPEG2000 EncoderabstractA high-speed VLSI design of an area-efficient JPEG2000 encoder is given. Recursive multilevel 2D discrete wavelet transform (DWT) architecture with dual buffers is proposed to reduce the wavelet coefficients memory to 1/4 tile size, prerate allocation is used to reduce the compressed code memory to 3/4 tile size. A highly pipelined and parallelism implementation of line-based 1-level DWT is proposed using two line-buffers in 5/3 wavelet type and its input speed is up to 2 samples/cycle; code block based address mapping in access wavelet coefficients memory, concurrent state variables generation and multiple parallel and pipeline coding methods are used in the bit plane encoder (BPE) which encodes on average at 40.5 M samples/s at 100 MHz with no memory used; the conditional two-symbol pipeline arithmetic encoder (AE) encodes at 1.3 symbols/cycle. Parallel units in BPE and buffer control between BPE and AE are optimally implemented with low cost without performance loss. Byte representation of rate-distortion slope used reaches a near optimal implementation of post-coding rate distortion in Tier2 with low cost. The compressed file generated by the encoder is fully compatible with ISO/IEC FCD15444-1. The encoder is verified on field-programmable gate array platform with a direct interface to digital video input with tile size 256 times 256 and code block size 32 times 16. The resulting input sampling rate is up to 58 M samples/s when Tier1 operates at 100 MHz. Difference of the peak signal-to-noise ratio of images compressed by our encoder and JasPer is less than 0.2 dB when the compression ratio is greater than 1 bps. Equivalent NAND2 gates synthesized are 90.6 K and on-chip RAM size is 626.75 kb. Unlike other designs the proposed design of JPEG2000 encoder has high compression quality as well as high speed and area-efficiency. Nanning Zheng 0001, Chang Huang, Yuehu Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2006 | Service Matchmaking with Rough SetsabstractWith the wide adoption of open grid services architecture (OGSA) and Web services resource framework (WSRF), the grid is emerging as a service-oriented computing infrastructure for engineers and scientists to solve data and computationally intensive problems. It is envisioned that computing resources in a future grid environment will be exposed as services. Service discovery becomes an issue of vital importance for a wider uptake of the grid. This paper presents RSSM, a rough sets based service matchmaking algorithm for service discovery with an aim to tolerate uncertainty in identifying service properties. The evaluation results show that the RSSM algorithm is more effective in service discovery compared with other mechanisms such as UDDI and OWLS. Maozhen Li 0001, Bin Yu 0005, Chang Huang, Yong-Hua Song |
CCGRID | 3 |
| 2005 | Vector Boosting for Rotation Invariant Multi-View Face DetectionabstractIn this paper, we propose a novel tree-structured multiview face detector (MVFD), which adopts the coarse-to-fine strategy to divide the entire face space into smaller and smaller subspaces. For this purpose, a newly extended boosting algorithm named vector boosting is developed to train the predictors for the branching nodes of the tree that have multicomponents outputs as vectors. Our MVFD covers a large range of the face space, say, +/-45/spl deg/ rotation in plane (RIP) and +/-90/spl deg/ rotation off plane (ROP), and achieves high accuracy and amazing speed (about 40 ms per frame on a 320 /spl times/ 240 video sequence) compared with previous published works. As a result, by simply rotating the detector 90/spl deg/, 180/spl deg/ and 270/spl deg/, a rotation invariant (360/spl deg/ RIP) MVFD is implemented that achieves real time performance (11 fps on a 320 /spl times/ 240 video sequence) with high accuracy. Chang Huang, Haizhou Ai, Yuan Li 0022, Shihong Lao |
ICCV | 1 |
| 2005 | Robust face alignment based on local texture classifiersabstractWe propose a robust face alignment algorithm with a novel discriminative local texture model. Different from the conventional descriptive PCA local texture model in ASM, classifiers using LUT-type Haar-like features are trained from a large data set as local texture model. The strong discriminative power of the classifier greatly improves the accuracy and robustness of local searching on faces with expression variation and ambiguous contours. A Bayesian framework is configured for shape parameter optimization and the algorithm is implemented in a hierarchical structure for efficiency. Extensive experiments are reported to show its accuracy and robustness. Haizhou Ai, Shengjun Xin, Chang Huang, Shuichiro Tsukiji, Shihong Lao |
ICIP (2) | 4 |
| 2004 | Exploiting multiple modalities for interactive video retrievalabstractAural and visual cues can be automatically extracted from video and used to index its contents. The paper explores the relative merits of the cues extracted from the different modalities for locating relevant shots in video, specifically reporting on the indexing and interface strategies used to retrieve information from the Video TREC 2002 and 2003 data sets, and the evaluation of the interactive search runs. For the documentary and news material in these sets, automated speech recognition produces rich textual descriptions derived from the narrative, with visual descriptions and depictions offering additional browsing functionality. Through speech and visual processing, storyboard interfaces with query-based filtering provide an effective interactive retrieval interface. Examples drawn from the Video TREC 2002 and 2003 search topics and results using these topics illustrate the utility of multiple-document storyboards and other interfaces incorporating the results of multimodal processing. Michael G. Christel, Chang Huang, Neema Moraveji, Norman Papernick |
ICASSP (3) | 2 |
| 2004 | Omni-directional face detection based on real adaboostabstractWe propose an omni-directional face detection method based on the confidence-rated AdaBoost algorithm, called real AdaBoost, proposed by R.E. Schapire and Y. Singer (see Machine Learning, vol.37, p.297-336, 1999). To use real AdaBoost, we configure the confidence-rated look-up-table (LUT) weak classifiers based on Haar-type features. A nesting-structured framework is developed to combine a series of boosted classifiers into an efficient object detector. For omni-directional face detection, our method has achieved a rather high performance and the processing speed can reach 217 ms per 320/spl times/240 image. Experiment results on the CMU+MIT frontal and the CMU profile face test sets are reported to show its effectiveness. Chang Huang, Bo Wu 0001, Haizhou Ai, Shihong Lao |
ICIP | 1 |
| 2004 | Evaluating content-based filters for image and video retrievalabstractThis paper investigates the level of metadata accuracy required for image filters to be valuable to users. Access to large digital image and video collections is hampered by ambiguous and incomplete metadata attributed to imagery. Though improvements are constantly made in the automatic derivation of semantic feature concepts such as indoor, outdoor, face, and cityscape, it is unclear how good these improvements should be and under what circumstances they are effective. This paper explores the relationship between metadata accuracy and effectiveness of retrieval using an amateur photo collection, documentary video, and news video. The accuracy of the feature classification is varied from performance typical of automated classifications today to ideal performance taken from manually generated truth data. Results establish an accuracy threshold at which semantic features can be useful, and empirically quantify the collection size when filtering first shows its effectiveness. Michael G. Christel, Neema Moraveji, Chang Huang |
SIGIR | 3 |
| 2003 | Enhanced access to digital video through visually rich interfacesabstractAn image-rich interface is presented, which emphasizes visual exploration of sets of images representing shots returned from a query of filter against a digital video corpus. This interface, a storyboard of keyframes for multiple video segments, maintains temporal layout, accommodates contextual cues and filtering, supports additional filtering through visual features, and provides a means of drilling down to synchronized points in the associated video. These features allow for effective information retrieval from a video collection, as evidenced by the success achieved in interactive query for the TREC 2002 Video Retrieval Track (TREC-V). This paper introduces TREC-V, discusses the design of the multi-segment storyboard interface, illustrates its use with respect to the TREC-V topics, and presents results and conclusions based on the TREC-V evaluation. Michael G. Christel, Chang Huang |
ICME | 2 |
| 2003 | A Negotiation Protocol for Database Resource BindingabstractThe negotiation process is a necessary precondition for many wide-area based applications which involve interactions across multiple autonomous entities. In the case where database resources are explicitly exposed to the distributed computing context like the Grid, it is necessary for unauthorized database users to get authorized so that they can be granted with certain privileges to access the desirable database resources. The facility to establish a legal identity for database user is one of the functionalities that the database resource management framework should provide. In this paper, we give a handshaking protocol which is applied between database clients and database service providers to achieve the database-specific negotiation. It's a six-phased protocol, which is described in terms of service operations in our database resource management middleware, which is an OGSA compliant service-based system. Chang Huang, Zhaohui Wu 0001, Guozhou Zheng |
ISPDC | 1 |
| 2003 | Open grid services of traditional Chinese medicineabstractA grid architecture, called TCM-grid for traditional Chinese medicine, is presented. Based on the open grid service architecture (OGSA), we develop a set of database grid services for discovering and accessing TCM database resources remotely and transparently. Upon these database grid services, a set of knowledge services are also specified to support TCM knowledge sharing in the Web environment. With our application experience, we argue that OGSA be promoted to support finely granular data sharing and knowledge-intensive tasks. Huajun Chen, Zhaohui Wu 0001, Chang Huang |
SMC | 3 |
| 2001 | SVG for navigating digital news videoabstractScalable Vector Graphics (SVG) is a language for describing two-dimensional graphics in XML, specifically vector graphic shapes, images, and text. SVG is a new World Wide Web Consortium (W3C) Candidate Recommendation as of November 2000, and this paper describes how SVG provides an ideal framework for presenting manipulable, interactive summarizations into a multimedia information repository. Specifically, we present VIBE and map SVG interfaces into a digital news video library for delivery through web browsers. Pan-and-zoom visualizations of video through SVG are discussed. Michael G. Christel, Chang Huang |
ACM Multimedia | 2 |