EDBT 2026 Demo / reviewers in the wild / expert
Luhong Liang
dblp:19/4685
· DBLP profile ↗
42ranked-venue papers
6as first author
10since 2021 · last 2026
0009-0001-4190-006XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 4 first-authorSystems, architecture and hardware · 14 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 6Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Adaptive Depth Processing System Based on Foundation Transformers and Time-of-Flight FusionabstractGenerating high-quality depth maps with accurate values is a critical research topic, and various depth estimation methods, such as Time-of-Flight (ToF) and Monocular Depth Estimation (MDE), are advancing rapidly. However, each single-sensor approach is inherently limited by the characteristics of its respective modality. Meanwhile, fused approaches using multiple sensors or dedicated trained models are often plagued by system complexity and limited generalization. In this paper, we propose an adaptive depth processing system based on the foundation transformer and ToF fusion, aiming to harness the precision of ToF data and the high-quality depth priors provided by the foundation model in a monocular and zero-shot manner. A hardware–algorithm co-design partitions the computation between a dedicated ASIC for compute-intensive foundation transformer acceleration and an MPSoC FPGA for potential reconfigurable fusion, yielding 36.4 ms latency and 1.9 W power under a 27.6 GOPS/frame workload. Zero-shot evaluations on various ToF data sources including FLAT, TICaM, and in-house captures confirm strong cross-domain generalization, demonstrating the system’s adaptability, efficiency, and accuracy. An-Nan Xiong, Pingcheng Dong, Yonghao Tan, Yuzhong Jiao, Manto Yung, Ann Li, Luhong Liang, Mansun Chan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2025 | A 10.60 μW 150 GOPS Mixed-Bit-Width Sparse CNN Accelerator for Life-Threatening Ventricular Arrhythmia DetectionabstractThis paper proposes an ultra-low power, mixed-bit-width sparse convolutional neural network (CNN) accelerator to accelerate ventricular arrhythmia (VA) detection. The chip achieves 50% sparsity in a quantized 1D CNN using a sparse processing element (SPE) architecture. Measurement on the prototype chip TSMC 40nm CMOS low-power (LP) process for the VA classification task demonstrates that it consumes 10.60 μW of power while achieving a performance of 150 GOPS and a diagnostic accuracy of 99.95%. The computation power density is only 0.57 μW/mm2, which is 14.23× smaller than state-of-the-art works, making it highly suitable for implantable and wearable medical devices. Zhenge Jia, Zheyu Yan, Jay Mok, Manto Yung, Yu Liu 0007, Wujie Wen, Luhong Liang, Kwang-Ting Cheng, Xiaobo Sharon Hu, Yiyu Shi 0001 |
ASP-DAC | 9 |
| 2025 | APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-DesignabstractDNN accelerators, significantly advanced by model compression and specialized dataflow techniques, have marked considerable progress. However, the frequent access of highprecision partial sums (PSUMs) leads to excessive memory demands in architectures utilizing input/weight stationary dataflows. Traditional compression strategies have typically overlooked PSUM quantization, which may account for 69% of power consumption. This study introduces a novel Additive Partial Sum Quantization (APSQ) method, seamlessly integrating PSUM accumulation into the quantization framework. A grouping strategy that combines APSQ with PSUM quantization enhanced by a reconfigurable architecture is further proposed. The APSQ performs nearly lossless on NLP and CV tasks across BERT, Segformer, and EfficientViT models while compressing PSUMs to INT8. This leads to a notable reduction in energy costs by $\mathbf{2 8-8 7 \%}$. Extended experiments on LLaMA2-7B demonstrate the potential of APSQ for large language models. Code is available at https://github.com/Yonghao-Tan/APSQ. Yonghao Tan, Pingcheng Dong, Yongkun Wu, Yu Liu 0007, Shih-Yang Liu, Xijie Huang, Luhong Liang, Kwang-Ting Cheng |
DAC | 10 |
| 2025 | SeDA: Secure and Efficient DNN Accelerators with Hardware/Software SynergyabstractEnsuring the confidentiality and integrity of DNN accelerators is paramount across various scenarios spanning autonomous driving, healthcare, and finance. However, current security approaches typically require extensive hardware resources, and incur significant off-chip memory access overheads. This paper introduces SeDA, which utilizes 1) a bandwidth-aware encryption mechanism to improve hardware resource efficiency, 2) optimal block granularity through intra-layer and inter-layer tiling patterns, and 3) a multi-level integrity verification mechanism that minimizes, or even eliminates, memory access overheads. Experimental results show that SeDA decreases performance overhead by over 12% for both server and edge neural processing units (NPUs), while ensuring robust scalability.11SeDA source code:https://github.com/wayne4s/seda.git Lang Feng 0001, Ning Lin, Zihao Xuan, Rongliang Fu, Tsung-Yi Ho, Yuzhong Jiao, Luhong Liang |
DAC | 9 |
| 2025 | CoXplorer: Multi-Staged Co-Exploration Framework for AI Model Compression and Accelerator DesignabstractThe rapid evolution of artificial intelligence (AI) algorithms demands efficient computing chips, positioning algorithm-hardware co-design as a crucial optimization strategy. However, automating the co-design process remains challenging due to the lack of a unified exploration framework for both algorithmic and hardware domains, as existing tools - hardware design space exploration (DSE) and compression neural architecture search (Compression NAS) - operate independently, relying entirely on manual collaboration. This paper presents CoXplorer, a co-exploration framework that connects model-compression optimization space and architecture design space. We make three key contributions: (1) a multi-staged co-design space decomposition method that enables systematic exploration of compression-hardware design choices with reduced complexity, (2) an AC-Copilot toolchain enhanced with multi-grained performance modeling driven by hardware simulation-compilation hierarchical cooperation to fulfill various evaluation requirements of co-exploration, enabling balanced simulation accuracy-efficiency trade-offs, and (3) a co-exploration workflow with hierarchical and bottleneck-guided search for harmonizing optimization objectives of both model and hardware design spaces, resulting in improved search efficiency. We validate the CoXplorer on two edge chips, which achieve 53.7% throughput and 45.8% energy efficiency improvements for the CNN acceleration, and 7.5× speedup with 9.9× energy efficiency boost for the Transformer acceleration. A case study on large language model acceleration shows CoXplorer’s extensibility to emerging workloads, enhancing LLAMA2-7B inference throughput from 6.75 to 25.46 tokens/s via co-optimization with compression and near-memory computing architecture. Songchen Ma, Yonghao Tan, Pingcheng Dong, Di Pang, Yu Liu 0007, Luhong Liang, Kwang-Ting Cheng, Fengbin Tu |
ICCAD | 9 |
| 2024 | Genetic Quantization-Aware Approximation for Non-Linear Operations in TransformersabstractNon-linear functions are prevalent in Transformers and their lightweight variants, incurring substantial and frequently underestimated hardware costs. Previous state-of-the-art works optimize these operations by piece-wise linear approximation and store the parameters in look-up tables (LUT), but most of them require unfriendly high-precision arithmetics such as FP/INT 32 and lack consideration of integer-only INT quantization. This paper proposed a genetic LUT-Approximation algorithm namely GQA-LUT that can automatically determine the parameters with quantization awareness. The results demonstrate that GQA-LUT achieves negligible degradation on the challenging semantic segmentation task for both vanilla and linear Transformer models. Besides, proposed GQA-LUT enables the employment of INT8-based LUT-Approximation that achieves an area savings of 81.3~81.7% and a power reduction of 79.3~80.2% compared to the high-precision FP/INT 32 alternatives. Code is available at https://github.com/PingchengDong/GQA-LUT. Pingcheng Dong, Yonghao Tan, Tianwei Ni, Yu Liu 0007, Luhong Liang, Shih-Yang Liu, Xijie Huang, Huaiyu Zhu 0004, Fengwei An, Kwang-Ting Cheng |
DAC | 8 |
| 2024 | An End-to-End Deep-Learning-Based Indirect Time-of-Flight Image Signal ProcessorabstractIndirect time-of-flight (iToF) is one of the most straightforward approaches to capture 3D images. However, due to the nature of the iToF camera, it is still challenging to get reliable and accurate depth images from the raw data in the image signal processor (ISP) pipeline due to environmental issues. Previous iToF ISP works mainly focus on the traditional pipeline, such as filters and depth calculation. In this work, we present an end-to-end deep-learning-based iToF ISP. The proposed iToF ISP system can generate real-time depth images with deep-learning-based noise reduction and multipath interference (MPI) reduction. With the mixed-bit convolutional neural network (CNN) with 96.5 % sparsity and the mixed-bit sparse accelerator, the CNN is accelerated by 2.78× and negligible mean average error (MAE) loss has been achieved on the FLAT dataset using the proposed ISP pipeline. An-Nan Xiong, Yuzhong Jiao, Manto Yung, Luhong Liang, Mansun Chan |
ISCAS | 6 |
| 2024 | WASP: Efficient Power Management Enabling Workload-Aware, Self-Powered AIoT DevicesabstractThe wide adoption of edge AI has heightened the demand for various battery-less and maintenance-free smart systems. Nevertheless, emerging Artificial Intelligence of Things (AIoT) are complex workloads showing increased power demand, diversified power usage patterns, and unique sensitivity to power management (PM) approaches. Existing AIoT devices cannot select the most appropriate PM tuning knob, and therefore they often make sub-optimal decisions. In addition, these PM solutions always assume traditional power regulation circuit which incurs non-negligible power loss and control overhead. This can greatly compromise the potential of AIoT efficiency. In this paper, we explore power management optimization for emerging self-powered AIoT devices. We propose WASP, a highly efficient power management scheme for workload-aware, self-powered AIoT devices. The novelty of WASP is two fold. First, it combines offline profiling and light-weight online control to select the most appropriate PM tuning knobs for the given DNN models. Second, it is well tailored to a reconfigurable voltage regulation module that can make the best use of the limited power budget. Our results show that WASP allows AIoT devices to accomplish 65.6% more inference tasks under a stringent power budget without any performance degradation compared with other existing approaches. Xiaofeng Hou, Xuehan Tang, Jiacheng Liu 0001, Chao Li 0009, Luhong Liang, Kwang-Ting Cheng |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2023 | Architecting Efficient Multi-modal AIoT SystemsabstractMulti-modal computing (M2C) has recently exhibited impressive accuracy improvements in numerous autonomous artificial intelligence of things (AIoT) systems. However, this accuracy gain is often tethered to an incredible increase in energy consumption. Particularly, various highly-developed modality sensors devour most of the energy budget, which would make the deployment of M2C for real-world AIoT applications a difficult challenge. Xiaofeng Hou, Jiacheng Liu 0001, Xuehan Tang, Chao Li 0009, Jia Chen 0032, Luhong Liang, Kwang-Ting Cheng, Minyi Guo |
ISCA | 6 |
| 2022 | ReAAP: A Reconfigurable and Algorithm-Oriented Array Processor With Compiler-Architecture Co-DesignabstractParallelism and data reuse are the most critical issues for the design of hardware acceleration in a deep learning processor. Besides, abundant on-chip memories and precise data management are intrinsic design requirements because most of deep learning algorithms are data-driven and memory-bound. In this paper, we propose a compiler-architecture co-design scheme targeting a reconfigurable and algorithm-oriented array processor, named ReAAP. Given specific deep neural networks, the proposed co-design scheme is effective to perform parallelism and data reuse optimization on compute-intensive layers for guiding reconfigurable computing in hardware. Especially, the systemic optimization is performed in our proposed domain-specific compiler to deal with the intrinsic tensions between parallelism and data locality, for the purpose of automatically mapping diverse layer-level workloads onto our proposed reconfigurable array architecture. In this architecture, abundant on-chip memories are software-controlled and its massive data access is precisely handled by compiler-generated instructions. In our experiments, the ReAAP is implemented on an embedded FPGA platform. Experimental results demonstrate that our proposed co-design scheme is effective to integrate software flexibility with hardware parallelism for accelerating diverse deep learning workloads. As a whole system, ReAAP achieves a consistently high utilization of hardware resource for accelerating all the diverse compute-intensive layers in ResNet, MobileNet, and BERT. Jianwei Zheng 0002, Yu Liu 0007, Luhong Liang, Deming Chen, Kwang-Ting Cheng |
IEEE Trans. Computers | 4 |
| 2020 | Optimize FPGA-Based Neural Network Accelerator with Bit-Shift QuantizationabstractWell-programmed Field Programmable Gate Arrays (FPGAs) can accelerate Deep Neural Network (DNN) with high power efficiency. The dominant workloads of DNNs are Multiply Accumulates (MACs), which can be directly mapped to Digital Signal Processors (DSPs) in the FPGA. A DNN accelerator pursuing high performance can consume almost all the DSPs, but with a considerable amount of Look-up Tables (LUTs) in the FPGA unused or performing MACs inefficiently. To solve this problem, we present a Bit-Shift method for FPGA-based DNN accelerator to fully utilize the resources in the FPGA. The MAC is converted to a limited number of shift-and-add operations, which can be implemented by LUTs with significant improvement of efficiency. A quantization method based on Minimum Mean Absolute Error (MMAE) is proposed to preserve the accuracy of the DNN inference in the conversion of DNN parameters without re-training. The quantized parameters can be compressed to a fixed and fewer number of bits to reduce the memory bandwidth. Accordingly, a Bit-Shift architecture is designed to load the compressed parameters and perform the converted MAC calculations without extra decompression module. A large scale DNN accelerator with the proposed Bit-Shift architecture is implemented in a Xilinx VU095 FPGA. Experimental results show that the proposed method can boost the processing speed by 32% and reach 331 GOPS at 190MHz clock frequency for ResNet-34. Yu Liu 0007, Luhong Liang |
ISCAS | 3 |
| 2017 | Retrain-free fully connected layer optimization using matrix factorizationabstractThe complexity of Deep Neural Networks (DNNs) hinders their implementation on embedded system with limited hardware resources. To deal with this issue, this paper presents a novel optimization algorithm based on semi-Nonnegative Matrix Factorization for Fully Connected layers (semi-NMF-based FC optimization). Compared with previous network surgery techniques, our proposed method optimizes network structure in a more implementation-friendly manner with full controllability and no requirement on retraining using training data. Simulations using AlexNet and VGG-16 on ImageNet task show the effectiveness and efficiency of semi-NMF-based FC optimization. Luhong Liang |
ICIP | 3 |
| 2016 | Deep learning neural networks optimization using hardware cost penaltyabstractAs deep learning neural networks (DNNs) advance and increase in computational complexity, particularly in terms of memory cost, it becomes difficult to implement DNNs in fixed-point memory-sparse environments (e.g. integrated circuits in consumer electronics). Thus, the training of DNNs must be reformulated to balance the hardware costs needed to represent the often millions of parameter weights in such machine learning models. This paper proposes a novel optimization approach that simultaneously minimizes complexity (total memory) and maximizes accuracy. Specifically, a bit-depth complexity penalty is induced to urge the DNN model towards a state of lower memory using a numerical gradient in optimization iterations. Experimental results on the MNIST (Mixed National Institute of Standards and Technology) handwritten digit classification benchmark demonstrate a minimal loss (1%) in DNN accuracy with a significant reduction (37%) of memory cost on average. Rohan Doshi, Kwok-Wai Hung, Luhong Liang, King Hung Chiu |
ISCAS | 3 |
| 2014 | DCT coefficients generation model for film grain noise and its application in super-resolutionabstractFilm grain noise (FGN) is generated by the procedure of capturing pictures using photographic film. Images with FGN are subjectively pleasing. However, FGN is difficult to compress and its pleasant features are difficult to preserve when the images are resized. So in literature, FGN is extracted first, then regenerated for the processed noise-free images. In this paper a new method is proposed to generate FGN. In contrast to some other models which generate FGN in spatial domain, our method captures the statistic feature of FGN in frequency domain. FGN is signal dependent and the signal independent scaled noise image (SNI) is obtained by scaling FGN by corresponding noise-free image. We model each discrete cosine transform (DCT) coefficient of SNI as a Gaussian random variable. The Gaussian model parameters can be estimated from the stack of all the blocks in SNI based on stationary assumption. Experimental results show that proposed model recovers FGN with similar visual and spectrum properties to the original FGN. We also apply proposed model in superresolution and the quality of resultant images are improved. Ting Sun 0001, Luhong Liang, King Hung Chiu, Pengfei Wan 0001, Oscar C. Au |
ICIP | 2 |
| 2014 | Fast single frame super-resolution using perceptual visibility optimizationabstractExample-based super-resolution (SR) approaches mostly reconstruct and optimize the high-resolution (HR) image according to objective criteria such as imaging model. However, in consumer electronics applications the ultimate goal is better subjective visual effect rather than higher objective texture visibility. In this paper, we propose a SR method using perceptual visibility optimization (PVO). A human visual system (HVS) preference model is built based on just noticeable distortion (JND) threshold that considers the structural regularity of the texture. Following the scale-invariant self-similarity (SiSS) based SR reconstruction, an iterative backprojection (IBP) procedure coupled with the proposed model adaptively enhances the texture visibility to match the HVS's preference. Experimental results show the proposed approach has competitive quality and lower computational complexity compared with several state-of-the-art SR approaches. Luhong Liang, Wai Keung Cheung, King Hung Chiu |
ISCAS | 1 |
| 2013 | Fast single frame super-resolution using scale-invariant self-similarityabstractExample-based super-resolution (SR) attracts great interest due to its wide range of applications. However, these algorithms usually involve patch search in a large database or the input image, which is computationally intensive. In this paper, we propose a scale-invariant self-similarity (SiSS) based super-resolution method. Instead of searching patches, we select the patch according to the SiSS measurement, so that the computational complexity is significantly reduced. Multi-shaped and multi-sized patches are used to collect sufficient patches for high-resolution (HR) image reconstruction and a hybrid weighting method is used to suppress the artifacts. Experimental results show that the proposed algorithm is 20~1,800 times faster than several state-of-the-art approaches and can achieve comparable quality. Luhong Liang, King Hung Chiu, Edmund Y. Lam |
ISCAS | 1 |
| 2013 | Strip Features for Fast Object DetectionabstractThis paper presents a set of effective and efficient features, namely strip features, for detecting objects in real-scene images. Although shapes of a specific class usually have large intraclass variance, some basic local shape elements are relatively stable. Based on this observation, we propose a set of strip features to describe the appearances of those shape elements. Strip features capture object shapes with edgelike and ridgelike strip patterns, which significantly enrich the efficient features such as Haar-like and edgelet features. The proposed features can be efficiently calculated via two kinds of approaches. Moreover, the proposed features can be extended to a perturbed version (namely, perturbed strip features) to alleviate the misalignment caused by deformations. We utilize strip features for object detection under an improved boosting framework, which adopts a complexity-aware criterion to balance the discriminability and efficiency for feature selection. We evaluate the proposed approach for object detection on the public data sets, and the experimental results show the effectiveness and efficiency of the proposed approach. Hong Chang 0001, Luhong Liang, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Cybern. | 3 |
| 2012 | A Fast and Performance-Maintained Transcoding Method Based on Background Modeling for Surveillance VideoabstractLow-complexity and high-performance surveillance video Transcoding methods play an important role for a wide range of surveillance video transmission and storage applications. Towards this end, the special characteristics of surveillance video should be utilized for Transcoding. In this paper, we propose a fast and performance-maintained Transcoding method. This method firstly divides macro blocks (MBs) into foreground MBs, foreground border MBs and background MBs. Statistics show that the three categories have different distributions of prediction modes, motion vectors and reference frames. Following this, we adopt different Transco ding strategies in terms of removing the redundant prediction modes, narrowing motion search range and reducing reference frames. In particular, we propose an algorithm to exploit the decoded motion vector to adaptively calculate motion search range. Experimental results show that, compared with the recent background modeling based full-decoding-full-encoding, our Transcoding method saves more than 93% time with ignorable quality loss. Mingchao Geng, Xianguo Zhang, Yonghong Tian 0001, Luhong Liang, Tiejun Huang 0001 |
ICME | 4 |
| 2012 | Macro-Block-Level Selective Background Difference Coding for Surveillance VideoabstractUtilizing the special properties to improve the surveillance video coding efficiency still has much room, although there have been three typical paradigms of methods: object-oriented, background-prediction-based and background-difference-based methods. However, due to the inaccurate foreground segmentation, the low-quality or unclear background frame, and the potential "foreground pollution" phenomenon, there is still much room for improvement. To address this problem, this paper proposes a macro-block-level selective background difference coding method (MSBDC). MSBDC selects the following two ways to encode each macro-block (MB): coding the original MB, and directly coding the difference data between the MB and its corresponding background. MSBDC also features at employs the classification of MBs to facilitate the selection, through which, prediction and motion compensation turns more accurate, both on foreground and background. Results show that, MSBDC significantly decreases the total bitrate and obtains a remarkable performance gain on foreground compared with several state-of-the-art methods. Xianguo Zhang, Yonghong Tian 0001, Luhong Liang, Tiejun Huang 0001, Wen Gao 0001 |
ICME | 3 |
| 2012 | Boosted translation-tolerable classifiers for fast object detection
Luhong Liang, Hong Chang 0001, Cherkeng Heng, Shiguang Shan, Xilin Chen 0001 |
Image Vis. Comput. | 2 |
| 2011 | A unified framework for locating and recognizing human actionsabstractIn this paper, we present a pose based approach for locating and recognizing human actions in videos. In our method, human poses are detected and represented based on deformable part model. To our knowledge, this is the first work on exploring the effectiveness of deformable part models in combining human detection and pose estimation into action recognition. Comparing with previous methods, ours have three main advantages. First, our method does not rely on any assumption on video preprocessing quality, such as satisfactory foreground segmentation or reliable tracking; Second, we propose a novel compact representation for human pose which works together with human detection and can well represent the spatial and temporal structures inside an action; Third, with human detection taken into consideration in our framework, our method has the ability to locate and recognize multiple actions in the same scene. Experiments on benchmark datasets and recorded cluttered videos verified the efficacy of our method. Yuelei Xie, Hong Chang 0001, Zhe Li 0008, Luhong Liang, Xilin Chen 0001, Debin Zhao |
CVPR | 4 |
| 2011 | GPU based fast algorithm for tanner graph based image interpolationabstractIn image/video processing software and hardware products, low complexity interpolation algorithms, such as cubic and splines methods, are commonly used. However, these methods tend to blur textures and produce jaggy effect compared with other adaptive methods such as NEDI, SAI. Tanner graph based image interpolation algorithm has better effect in dealing with edge and texture, but with high computation complexity. Thanks to the high performance parallel processing capability of today's GPU, use of complex algorithms for real time application is becoming possible. In this paper, we present a fast algorithm for tanner graph based image interpolation and it's implementation on GPU. In our algorithm, the image model training process of tanner graph based image interpolation is greatly simplified. Experimental results show that the GPU implementation can be more than 47 times as fast as the CPU implementation. Wei Lei, Ruiqin Xiong, Siwei Ma 0001, Luhong Liang |
MMSP | 4 |
| 2011 | Fast intra mode selection for stereo video coding using epipolar constraintabstractApplication of stereo video in TV industry and consumer electronics becomes very popular recently. Thus, fast algorithm for stereo video coding is highly desired because of its huge inter-view computational redundancy. In our previous work we proposed an epipolar constraint based fast inter mode selection for stereo video using motion vector of blocks as a indicator of similarity, nevertheless, intra mode selection is also highly complex in standards like H.264/AVC. In this paper, a practical fast intra mode selection is proposed to eliminate the computational redundancy by exploiting inter-view dependency based on epipolar constraint. The proposed method does not rely on disparity estimation. Instead, a sliding window is employed to generate an intra mode candidate pool from macro-blocks on the epipolar line. The candidate pool is then rectified to remove invalid modes and improve accuracy. Finally, optimal prediction mode is selected from the candidate pool. The proposed method can significantly reduce the number of mode candidates/prediction directions compared to exhaustive mode selection by 79%. Experiments on 5 HD video coded in 1-frame demonstrate the overall coding time of one view is saved by 56% on average, with slightly video quality loss less than 0.1 dB. Guolei Yang, Luhong Liang, Wen Gao 0001 |
VCIP | 2 |
| 2010 | Fast object detection using boosted co-occurrence histograms of oriented gradientsabstractCo-occurrence histograms of oriented gradients (CoHOG) are powerful descriptors in object detection. In this paper, we propose to utilize a very large pool of CoHOG features with variable-location and variable-size blocks to capture salient characteristics of the object structure. We consider a CoHOG feature as a block with a special pattern described by the offset. A boosting algorithm is further introduced to select the appropriate locations and offsets to construct an efficient and accurate cascade classifier. Experimental results on public datasets show that our approach simultaneously achieves high accuracy and fast speed on both pedestrian detection and car detection tasks. Cherkeng Heng, Luhong Liang, Xilin Chen 0001 |
ICIP | 4 |
| 2010 | A Sample Pre-mapping Method Enhancing Boosting for Object DetectionabstractWe propose a novel method to improve the training efficiency and accuracy of boosted classifiers for object detection. The key step of the proposed method is a sample pre-mapping on original space by referring to the selected `reference sample' before feeding into weak classifiers. The reference sample corresponds to an approximation of the optimal separating hyper-plane in an implicit high dimensional space, so that the resulting classifier could achieve the performance similar to kernel method, while spending the computation cost of linear classifier in both training and detection. We employ two different non-linear mappings to verify the proposed method under boosting framework. Experimental results show that the proposed approach achieves performance comparable with the common used methods on public datasets in both pedestrian detection and car detection. Xiaopeng Hong, Cherkeng Heng, Luhong Liang, Xilin Chen 0001 |
ICPR | 4 |
| 2010 | An epipolar resticted inter-mode selection for stereoscopic video encodingabstractFast stereoscopic video encoding becomes a highly desired technique because the stereoscopic video has been realizable for applications like TV broadcasting and consumer electronics. The stereoscopic video has high inter-view dependency subject to epipolar restriction, which can be used to reduce the encoding complexity. In this paper, we propose a fast inter-prediction mode selection algorithm for stereoscopic video encoding. Different from methods using disparity estimation, candidate modes are generated by sliding a window along the macro-block line restricted by the epipolar. Then the motion information is utilized to rectify the candidate modes. A selection failure handling algorithm is also proposed to preserve coding quality. The proposed algorithm is evaluated using independent H.264/AVC encoders for left and right views and can be extended to MVC. Experimental results show that encoding times of one view are reduced by 41.4% and 24.4% for HD and VGA videos respectively with little quality loss. Guolei Yang, Luhong Liang, Wen Gao 0001 |
PCS | 2 |
| 2010 | A background model based method for transcoding surveillance videos captured by stationary cameraabstractReal-world video surveillance applications require storing videos without neglecting any part of scenarios for weeks or months. To reduce the storage cost, the high bit-rate videos from cameras should be transcoded into a more efficient compressed format with as little quality loss as possible. In this paper, we propose a background model based method to improve the transcoding efficiency for surveillance videos captured by stationary cameras, and objectively measure it. The background model is trained by pre-decoded I frames, and then used to transcode the source stream. Following this method, an H.264/AVC based transcoder employing the background model as long-term reference frame and a difference frame coding based transcoder are implemented and evaluated. Experimental results show that both trancoders save nearly half the used bits while maintaining quality compared with the full-decoding-full-encoding method, and the latter one has slightly better performance. Xianguo Zhang, Luhong Liang, Qian Huang 0008, Tiejun Huang 0001, Wen Gao 0001 |
PCS | 2 |
| 2010 | A fast intra 4×4 mode decision algorithm for H.264/AVC down rate transcodingabstractH.264/AVC adopts 9 intra 4×4 prediction modes to improve the intra frame compression efficiency. This makes the intra frame coding and transcoding very complex. In this paper, a fast intra 4×4 mode decision algorithm is proposed to reduce the intra transcoding complexity. Firstly, the surrounding modes in the high bitrate and the modes of the neighboring blocks in the current bitrate are used to form the candidate modes. Secondly, the variance and mean of the intra prediction residual are used to further reduce the size of the candidate mode set. Thirdly, the intra mode context model is built from the high bitrate modes and updated by the low bitrate modes to order the candidate modes. An early skip scheme is used to terminate the mode decision process. Experimental results show that about 50% intra 4×4 modes can be saved with a penalty on coding efficiency less than 0.1dB. The transcoding performance can even gain up to 0.6 dB on some sequences. Zhihang Wang, Luhong Liang, Shengfu Dong, Wen Gao 0001, Debin Zhao, Qingming Huang |
VCIP | 2 |
| 2010 | An efficient coding scheme for surveillance videos captured by stationary camerasabstractIn this paper, a new scheme is presented to improve the coding efficiency of sequences captured by stationary cameras (or namely, static cameras) for video surveillance applications. We introduce two novel kinds of frames (namely background frame and difference frame) for input frames to represent the foreground/background without object detection, tracking or segmentation. The background frame is built using a background modeling procedure and periodically updated while encoding. The difference frame is calculated using the input frame and the background frame. A sequence structure is proposed to generate high quality background frames and efficiently code difference frames without delay, and then surveillance videos can be easily compressed by encoding the background frames and difference frames in a traditional manner. In practice, the H.264/AVC encoder JM 16.0 is employed as a build-in coding module to encode those frames. Experimental results on eight in-door and out-door surveillance videos show that the proposed scheme achieves 0.12 dB~1.53 dB gain in PSNR over the JM 16.0 anchor specially configured for surveillance videos. Xianguo Zhang, Luhong Liang, Qian Huang 0008, Yazhou Liu, Tiejun Huang 0001, Wen Gao 0001 |
VCIP | 2 |
| 2010 | No-reference perceptual image quality metric using gradient profiles for JPEG2000
Luhong Liang, Shiqi Wang 0001, Siwei Ma 0001, Debin Zhao, Wen Gao 0001 |
Signal Process. Image Commun. | 1 |
| 2009 | Fast car detection using image strip featuresabstractThis paper presents a fast method for detecting multi-view cars in real-world scenes. Cars are artificial objects with various appearance changes, but they have relatively consistent characteristics in structure that consist of some basic local elements. Inspired by this, we propose a novel set of image strip features to describe the appearances of those elements. The new features represent various types of lines and arcs with edge-like and ridge-like strip patterns, which significantly enrich the simple features such as haar-like features and edgelet features. They can also be calculated efficiently using the integral image. Moreover, we develop a new complexity-aware criterion for RealBoost algorithm to balance the discriminative capability and efficiency of the selected features. The experimental results on widely used single view and multi-view car datasets show that our approach is fast and has good performance. Luhong Liang |
CVPR | 2 |
| 2009 | A no-reference perceptual blur metric using histogram of gradient profile sharpnessabstractNo-reference measurement of blurring artifacts in images is a challenging problem in image quality assessment field. One of the difficulties is that the inherently blurry regions in some natural images may disturb the evaluation of blurring artifacts. In this paper, we study the image gradients along local image structures and propose a new perceptual blur metric to deal with the above problem. The gradient profile sharpness of image edge is efficiently calculated along horizontal or vertical direction. Then the sharpness distribution histogram rectified by just noticeable distortion (JND) threshold is used to evaluate the blurring artifacts and assess the image quality. Experimental results show that the proposed method can achieve good image quality prediction performance. Luhong Liang, Siwei Ma 0001, Debin Zhao, Wen Gao 0001 |
ICIP | 1 |
| 2004 | A multi-stream audio-video large-vocabulary Mandarin Chinese speech databaseabstractWe present the acquisition and content of a multi-stream audio-visual large-vocabulary database in Mandarin Chinese. The database consists of 17,000 utterances spoken by 225 people and captured by a set of seven cameras and 12 microphones. We also provide the label files that describe the endpoints of the utterances and the script files that represent the actual pronunciation of speech. The database can be used in audio-visual speech recognition (AVSR) for both large-vocabulary and small tasks, microphone array based speech recognition, audio-visual speaker identification and 3D face modeling. Luhong Liang, Feiyue Huang, Ara V. Nefian |
ICME | 1 |
| 2003 | Environment-adaptive multi-channel biometricsabstractThe paper looks into multi-channel/multimodal biometric systems that are adaptive to environmental variations. We introduce a general formulation that addresses the environmental robustness of multi-channel fusion in biometric systems. Based on the formulation, two audio-visual biometric systems are developed. The first relies on confidence measures derived from environmental conditions to weight the contributions of the biometric channels dynamically; whereas the second considers the multiple channels jointly to adjust the fusion parameters optimally according to the current environmental conditions. Experimental evaluations with varying testing conditions show that both systems achieve lower recognition error rate compared with a baseline non-environment-adaptive audio-visual system. It is further shown that incorporating joint optimization of multichannel fusion parameters to cater to environmental changes, as in the second system, consistently leads to improved recognition accuracy over other systems, and, at the same time, it is guaranteed to perform no worse than any of the individual biometric channels under all environmental conditions. Stephen M. Chu, Minerva M. Yeung, Luhong Liang, Xiaoxing Liu |
ICASSP (5) | 3 |
| 2003 | Audio-visual speaker identification using coupled hidden Markov modelsabstractIn this paper, we investigate the use of the coupled hidden Markov models (CHMM) for the task of audio-visual text dependent speaker identification. Our system determines the identity of the user from a temporal sequence of audio and visual observations obtained from the acoustic speech and the shape of the mouth, respectively. The multi modal observation sequences are then modeled using a set of CHMMs, one for each phoneme-viseme pair and for each person in the database. The use of CHMMs in our system is justified by the capacity of this model to describe the natural audio and visual state asynchrony as well as their conditional dependency over time. To train a CHMM we first train a speaker independent model using expectation-maximization (EM), and then we build a speaker dependent model using maximum a posteriori (MAP) training. Experimental results on XM2VTS database show that our system improves the accuracy of audio-only or video-only speaker identification at all levels of acoustic signal-to-noise ratio (SNR) from 0 to 30 dB. Tieyan Fu, Xiaoxing Liu, Luhong Liang, Xiaobo Pi, Ara V. Nefian |
ICIP (3) | 3 |
| 2003 | A detector tree of boosted classifiers for real-time object detection and trackingabstractThis paper presents a novel tree classifier for complex object detection tasks together with a general framework for real-time object tracking in videos using the novel tree classifier. A boosted training algorithm with a clustering-and-splitting step is employed to construct branches in the nodes recursively, if and only if it improves the discriminative power compared to a single monolithic node classifier and has a lower computational complexity. A mouth tracking system that integrates the tree classifier under the proposed framework is built and tested on XM2FDB database. Experimental results show that the detection accuracy is equal or better than a single or multiple cascade classifier, while being computational less demanding. Rainer Lienhart, Luhong Liang, Alexander Kuranov |
ICME | 2 |
| 2003 | Multi-Cue-Based Face and Facial Feature Detection on Video Segments
Zhenyun Peng, Haizhou Ai, Luhong Liang, Guangyou Xu |
J. Comput. Sci. Technol. | 4 |
| 2002 | A coupled HMM for audio-visual speech recognitionabstractIn recent years several speech recognition systems that use visual together with audio information showed significant increase in performance over the standard speech recognition systems. The use of visual features is justified by both the bimodality of the speech generation and by the need of features that are invariant to acoustic noise perturbation. The audio-visual speech recognition system presented in this paper introduces a novel audio-visual fusion technique that uses a coupled hidden Markov model (HMM). The statistical properties of the coupled-HMM allow us to model the state asynchrony of the audio and visual observations sequences while still preserving their natural correlation over time. The experimental results show that the coupled HMM outperforms the multistream HMM in audio visual speech recognition. Ara V. Nefian, Luhong Liang, Xiaobo Pi, Xiaoxiang Liu, Crusoe Mao, Kevin Murphy 0002 |
ICASSP | 2 |
| 2002 | Speaker independent audio-visual continuous speech recognitionabstractThe increase in the number of multimedia applications that require robust speech recognition systems determined a large interest in the study of audio-visual speech recognition (AVSR) systems. The use of visual features in AVSR is justified by both the audio and visual modality of the speech generation and the need for features that are invariant to acoustic noise perturbation. The speaker independent audio-visual continuous speech recognition system presented relies on a robust set of visual features obtained from the accurate detection and tracking of the mouth region. Further, the visual and acoustic observation sequences are integrated using a coupled hidden Markov (CHMM) model. The statistical properties of the CHMM can model the audio and visual state asynchrony while preserving their natural correlation over time. The experimental results show that the current system tested on the XM2VTS database reduces by over 55% the error rate of the audio only speech recognition system at SNR of 0 dB. Luhong Liang, Xiaoxing Liu, Yibao Zhao, Xiaobo Pi, Ara V. Nefian |
ICME (2) | 1 |
| 2002 | Audio-visual continuous speech recognition using a coupled hidden Markov modelabstractWith the increase in the computational complexity of recent computers, audio-visual speech recognition (AVSR) became an attractive research topic that can lead to a robust solution for speech recognition in noisy environments. In the audio visual continuous speech recognition system presented in this paper, the audio and visual observation sequences are integrated using a coupled hidden Markov model (CHMM). The statistical properties of the CHMM can describe the asyncrony of the audio and visual features while preserv-ing their natural correlation over time. The experimental re-sults show that the current system tested on the XM2VTS database reduces the error rate of the audio only speech recognition system at SNR of 0db by over 55%. 1. Xiaoxing Liu, Yibao Zhao, Xiaobo Pi, Luhong Liang, Ara V. Nefian |
INTERSPEECH | 4 |
| 2001 | Face detection based on template matching and support vector machinesabstractA face detection algorithm integrating template matching and support vector machines (SVM) is presented. Two types of templates: eyes-in-whole and face itself, are used for coarse filtering, and the SVM classifier is used for classification. A bootstrap method is used to collect non-face samples for SVM training under a template matching constrained subspace, which greatly reduces the complexity of training the SVM. Comparative experimental results demonstrate its effectiveness. Haizhou Ai, Luhong Liang, Guangyou Xu |
ICIP (1) | 2 |
| 2000 | A General Framework for Face Detection
Haizhou Ai, Luhong Liang, Guangyou Xu |
ICMI | 2 |