VLDB 2026 Research / reviewers in the wild / expert
Hua Cai
dblp:01/682
· DBLP profile ↗
36ranked-venue papers
13as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-authorHuman-computer interaction and ubiquitous computing · 4 · 3 first-authorArtificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Computer networks · 3 · 1 first-authorSystems, architecture and hardware · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Upper bound of the list r-hued chromatic number
Hong-Jian Lai, Hua Cai |
Discret. Appl. Math. | 5 |
| 2025 | CocoER: Aligning Multi-Level Feature by Competition and Coordination for Emotion RecognitionabstractWith the explosion of human-machine interaction, emotion recognition has reignited attention. Previous works focus on improving visual feature fusion and reasoning from multiple image levels. Although it is non-trivial to deduce a person’s emotion by integrating multi-level feature (head, body and context), the emotion recognition results of each level is usually different from one another, which creates inconsistency in the prevailing feature alignment method and decrease recognition performance. In this work, we propose a multi-level image feature refinement method for emotion recognition (CocoER) to mitigate the impact caused by conflicting results from multi-level recognition. First, we leverage cross-level attention to improve visual feature consistency between hierarchically cropped head, body and context windows. Then, vocabulary informed alignment is incorporated into the recognition framework to produce pseudo label and guide hierarchical visual feature refinement. To effectively fuse multi-level feature, we elaborate on a competition process of eliminating irrelevant image level predictions and a coordination process to enhance the feature across all levels. Extensive experiments are executed on two popular datasets, and our method achieves state-of-the-art performance with multi-level interpretation results. Code is available at: https://github.com/bisno/CocoER. Xuli Shen, Hua Cai, Weilin Shen, Qing Xu 0017, Dingding Yu, Weifeng Ge, Xiangyang Xue 0001 |
CVPR | 2 |
| 2025 | Unilaw-R1: A Large Language Model for Legal Reasoning with Reinforcement Learning and Iterative InferenceabstractReasoning-focused large language models (LLMs) are rapidly evolving across various domains, yet their capabilities in handling complex legal problems remains underexplored.In this paper, we introduce Unilaw-R1, a large language model tailored for legal reasoning.With a lightweight 7-billion parameter scale, Unilaw-R1 significantly reduces deployment cost while effectively tackling three core challenges in the legal domain: insufficient legal knowledge, unreliable reasoning logic, and weak business generalization.To address these issues, we first construct Unilaw-R1-Data, a high-quality dataset containing ∼17K distilled and screened chain-of-thought (CoT) samples.Based on this, we adopt a two-stage training strategy combining Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL), which significantly boosts the model's performance on complex legal reasoning tasks and supports interpretable decision-making in legal AI applications.To assess legal reasoning ability, we also introduce Unilaw-R1-Eval, a dedicated benchmark designed to evaluate models across single-and multi-choice legal tasks.Unilaw-R1 demonstrates strong results on authoritative benchmarks, outperforming all models of similar scale and achieving performance on par with the much larger DeepSeek-R1-Distill-Qwen-32B (54.9%).Following domainspecific training, it also showed significant gains on LawBench and LexEval, exceeding Qwen-2.5-7B-Instruct(46.6%) by an average margin of 6.6%.Code is available at: https://github.com/Hanscal/Unilaw-R1. Hua Cai, Xuli Shen, Qing Xu 0017, Weilin Shen, Tianke Ban |
EMNLP | 1 |
| 2025 | Enhancing Multimodal Analogical Reasoning Through Triplet InteractionabstractAnalogical reasoning is fundamental to human cognition and plays a crucial role across various fields. However, previous studies have primarily focused on single-modal analogical reasoning, often ignoring the benefits of incorporating structural knowledge. Research in cognitive psychology has shown that information from multimodal sources brings more powerful cognitive transfer than single-modality sources. Motivated by the cognitive analogical theory, we address the challenge of multimodal analogical reasoning over multimodal knowledge graphs. Our approach introduces a novel triplet interaction mechanism, where both internal and external interactions between entities and relations are leveraged to enhance the accuracy of multimodal analogical reasoning. Experimental results show that our model outperforms Transformer-based multimodal pre-trained baselines. This improvement underscores the essential role of triplet interactions in integrating multimodal data, leading to more accurate analogical reasoning processes. Hua Cai, Xuli Shen, Weilin Shen, Qing Xu 0017 |
ICASSP | 1 |
| 2025 | EmoHead: Emotional Talking Head via Manipulating Semantic Expression ParametersabstractGenerating emotion-specific talking head videos from audio input is an important and complex challenge for human-machine interaction. However, emotion is highly abstract concept with ambiguous boundaries, and it necessitates disentangled expression parameters to generate emotionally expressive talking head videos. In this work, we present EmoHead to synthesize talking head videos via semantic expression parameters. To predict expression parameter for arbitrary audio input, we apply an audio-expression module that can be specified by an emotion tag. This module aims to enhance correlation from audio input across various emotions. Furthermore, we leverage pre-trained hyperplane to refine facial movements by probing along the vertical direction. Finally, the refined expression parameters regularize neural radiance fields and facilitate the emotion-consistent generation of talking head videos. Experimental results demonstrate that semantic expression parameters lead to better reconstruction quality and controllability. Xuli Shen, Hua Cai, Dingding Yu, Weilin Shen, Qing Xu 0017, Xiangyang Xue 0001 |
ICME | 2 |
| 2023 | I will only know after using it: The repeat purchasers of smart home appliances and the privacy paradox problem
Tianfeng Li, Hua Cai, Jian Zhang 0080 |
Comput. Secur. | 3 |
| 2022 | An Electromagnetic Information Methodology for Fast MIMO Deterministic Channel AnalysisabstractThe electromagnetic information theory (EIT) is a general methodology framework, which should be able to achieve the wireless network design and performance analysis. Its one of the main objective is exploring how to increase spectral efficiency of the multiple input multiple output (MIMO) communication. The array arrangements, mutual-coupling effects between antennas, and antenna radiation characteristics may play important roles. Therefore, the electromagnetic (EM) theory will be a non-negligible analysis tool in the EIT. In this paper, we present an electromagnetic information methodology (EIM) combining the EM theory with information theory for the fast and accurate MIMO channel analysis. This methodology decouples MIMO antenna arrays with propagation environments, of which advantages are as follows: 1) for channel analysis of variable MIMO array arrangements, it avoids inefficient recalculation of complex and large-scale propagation environments; 2) It can determine the antenna radiation characteristics that achieves capacity maximization for the given propagation scenario; 3) The EM computation techniques can be applied in this framework, which makes the methodology more practical. Xianjin Li, Guangjian Wang, Hua Cai, Jia He 0002, Ziming Yu |
VTC Spring | 3 |
| 2021 | Fangorn: Adaptive Execution Framework for Heterogeneous Workloads on Shared ClustersabstractPervasive needs for data explorations at all scales have populated modern distributed platforms with workloads of different characteristics. The growing complexities and diversities have thereafter imposed distinct challenges to execute them on shared clusters in corporate or public clouds. This paper presents Fangorn, an adaptive execution framework built on an enriched graph model. As the underlying infrastructure for core computation platforms at Alibaba, Fangorn supports various execution modes and caters to heterogeneous workloads. With the capability to orchestrate graph executions with both long-running and requested-on-demand resources at the same time, Fangorn allows exploration of tradeoffs between latency and resource efficiency, for jobs of all scales. By modeling distributed job executions as mutable graphs with pluggable components, Fangorn offers a systematic framework to adjust job executions adaptively, according to data statistics collected during run-time. Fangorn supports an array of different computation engines ranging from relational to deep learning, and is fully deployed on production clusters across Alibaba. It manages tens of millions of distributed jobs daily, with job size scaling from one to half-million. Yingda Chen, Jiamang Wang, Yifeng Lu, Zhiqiang Lv, Xuebin Min, Hua Cai, Wei Zhang 0012, Haochuan Fan, Chao Li 0009, Wei Lin 0016, Yangqing Jia, Jingren Zhou 0001 |
Proc. VLDB Endow. | 7 |
| 2017 | Single-Image Super-Resolution for Remote Sensing Data Using Deep Residual-Learning Neural Network
Ningbo Huang, Xinchao Gu, Hua Cai |
ICONIP (2) | 5 |
| 2017 | A novel multiple-frequency inverter topology for inductively coupled power transfer systemabstractHigh frequency is an effective method to improve transmission efficiency of inductively coupled power transfer (ICPT) system. However, it is hard for the high capacity semiconductor device to meet the high frequency need. In this paper, a novel multiple-frequency inductively power transfer system is proposed. By adopting a novel multiple-frequency topology of inverter, the transfer frequency of the system could be several times the inverter's switching frequency. A doubled-frequency half bridge inverter is proposed in the paper, and it is analyzed and compared with the traditional one. Simulation and experiment results indicate that doubled-frequency inverter based inductively power transfer system can improve the transfer frequency effectively and the transmission efficiency is increased especially for high power application. Hua Cai, Liming Shi |
IECON | 1 |
| 2013 | An online service-oriented performance profiling tool for cloud computing systems
Haibo Mi, Huaimin Wang 0001, Yangfan Zhou 0002, Michael R. Lyu, Hua Cai, Gang Yin |
Frontiers Comput. Sci. | 5 |
| 2013 | Toward Fine-Grained, Unsupervised, Scalable Performance Diagnosis for Production Cloud Computing SystemsabstractPerformance diagnosis is labor intensive in production cloud computing systems. Such systems typically face many real-world challenges, which the existing diagnosis techniques for such distributed systems cannot effectively solve. An efficient, unsupervised diagnosis tool for locating fine-grained performance anomalies is still lacking in production cloud computing systems. This paper proposes CloudDiag to bridge this gap. Combining a statistical technique and a fast matrix recovery algorithm, CloudDiag can efficiently pinpoint fine-grained causes of the performance problems, which does not require any domain-specific knowledge to the target system. CloudDiag has been applied in a practical production cloud computing systems to diagnose performance problems. We demonstrate the effectiveness of CloudDiag in three real-world case studies. Haibo Mi, Huaimin Wang 0001, Yangfan Zhou 0002, Michael R. Lyu, Hua Cai |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2012 | P-Tracer: Path-Based Performance Profiling in Cloud Computing SystemsabstractIn large-scale cloud computing systems, the growing scale and complexity of component interactions pose great challenges for operators to understand the characteristics of system performance. Performance profiling has long been proved to be an effective approach to performance analysis; however, existing approaches do not consider two new requirements that emerge in cloud computing systems. First, the efficiency of the profiling becomes of critical concern; second, visual analytics should be utilized to make profiling results more readable. To address the above two issues, in this paper, we present P-Tracer, an online performance profiling approach specifically tailored for large-scale cloud computing systems. P-Tracer constructs a specific search engine that adopts a proactive way to process performance logs and generates particular indices for fast queries; furthermore, PTracer provides users with a suite of web-based interfaces to query statistical information of all kinds of services, which helps them quickly and intuitively understand system behavior. The approach has been successfully applied in Alibaba Cloud Computing Inc. to conduct online performance profiling both in production clusters and test clusters. Experience with one real-world case demonstrates that P-Tracer can effectively and efficiently help users conduct performance profiling and localize the primary causes of performance anomalies. Haibo Mi, Huaimin Wang 0001, Hua Cai, Yangfan Zhou 0002, Michael R. Lyu, Zhenbang Chen 0001 |
COMPSAC | 3 |
| 2012 | Performance problems diagnosis in cloud computing systems by mining request trace logsabstractIn cloud computing systems, end-to-end request tracing approach is helpful for developers to understand the runtime behavior of user requests. Based on trace logs, we propose an approach to localize the abnormal methods that are the primary causes of performance problems. Our approach involves three steps: (1) cluster the user requests into different categories according to request call sequences and select major categories; (2) extract the principal methods that might be the causes of performance degradation; (3) pick out abnormal methods from those principal methods in each major category. We conduct four cases of performance degradations to validate our approach over a real-world enterprise-class cloud computing platform. The experimental results show that our approach can locate the prime causes of performance problems with low false-positive rate and false-negative rate. Haibo Mi, Huaimin Wang 0001, Gang Yin, Hua Cai, Tingtao Sun |
NOMS | 4 |
| 2012 | Localizing root causes of performance anomalies in cloud computing systems by analyzing request trace logs
Haibo Mi, Huaimin Wang 0001, Yangfan Zhou 0002, Michael R. Lyu, Hua Cai |
Sci. China Inf. Sci. | 5 |
| 2012 | Coordinating Cognitive Assistance With Cognitive Engagement Control Approaches in Human-Machine CollaborationabstractIn human–machine collaboration, automated machines may assist operators in a variety of ways. However, chaotic assistance may lead to negative consequences, which makes the achievement of effective coordination of the different types of assistance all the more important. This paper discusses the classification of assistance on a cognitive basis and a method of coordinating assistance. Cognitive assistance is viewed as a 2-D problem, consisting of when to provide assistance (a control problem) and what assistance to provide (an interface problem). This paper further proposes dynamically controlling cognitive engagement levels to meet the demands of maintaining performance. Cognitive engagement control determines the appropriate moment to provide the proper level of cognitive assistance. To validate the above approach, a driving assistance experiment was conducted on a driving simulator. In the experiment, an intelligent assistance system monitored the real-time driving performance of human drivers, e.g., time headway and lateral deviation. Because of the importance of visual attention in driving performance, the system monitored the cognitive engagement status of drivers by measuring their eye movements with an eye tracker. Through five sessions of car-following driving tests, the coordinated cognitive assistance (named adaptive assistance) was compared with four other types of cognitive assistance: no aid, soft aid , soft intervention, and hard intervention. The experimental results confirmed that coordinated cognitive assistance is the most effective approach to provide assistance in both primary and secondary tasks. It also appears to be more enjoyable and less intrusive when compared with other individual types of cognitive assistance. Hua Cai, Yingzi Lin |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 2011 | Advance in grey incidence analysis modellingabstractA systematic carding on the research of grey incidence analysis modeling has been made in this paper. The grey incidence analysis models developed from the models based on incidence coefficients of each point in the sequences in early days to the generalized grey incidence analysis models based on integral or overall perspective. It evolved from the grey incidence analysis models which measure similarity based on nearness into the models which consider similarity and nearness respectively. The objects of the research advanced from the analysis of relationship among curves to that among curved surfaces, and further to the analysis of relationship in three-dimensional space and even the relationship among super surfaces in n-dimensional space. The problems remained to be studied in this field are clarified too. Several research approaches of grey incidence analysis modeling are clearly revealed. Sifeng Liu, Hua Cai, Yingjie Yang |
SMC | 2 |
| 2011 | Modeling of operators' emotion and task performance in a virtual driving environment
Hua Cai, Yingzi Lin |
Int. J. Hum. Comput. Stud. | 1 |
| 2010 | Comparative study of discretization methods of microarray data for inferring transcriptional regulatory networksabstractBACKGROUND: Microarray data discretization is a basic preprocess for many algorithms of gene regulatory network inference. Some common discretization methods in informatics are used to discretize microarray data. Selection of the discretization method is often arbitrary and no systematic comparison of different discretization has been conducted, in the context of gene regulatory network inference from time series gene expression data. RESULTS: In this study, we propose a new discretization method "bikmeans", and compare its performance with four other widely-used discretization methods using different datasets, modeling algorithms and number of intervals. Sensitivities, specificities and total accuracies were calculated and statistical analysis was carried out. Bikmeans method always gave high total accuracies. CONCLUSIONS: Our results indicate that proper discretization methods can consistently improve gene regulatory network inference independent of network modeling algorithms and datasets. Our new method, bikmeans, resulted in significant better total accuracies than other methods. Yong Li 0012, Hua Cai, Dianjing Guo, Yanming Zhu 0002 |
BMC Bioinform. | 4 |
| 2009 | Measurement and Analysis of One-Way Delays over IEEE 802.16e/WiBro NetworkabstractOne-way delay is an important metric to evaluate the performance of communication networks. In this paper, we analyze the one-way delay performance of the commercial WiBro networks based on the traffic measurement. In measurements, much large delay is observed in the uplink than downlink, and we find that such asymmetry due mainly to the impact of the uplink bandwidth request mechanism. Due to this large uplink delay, the performance of commercial IEEE 802.16e service can be severely degraded. We also discuss some possible solutions to ameliorate this problem. Dongmyoung Kim, Hua Cai, Sunghyun Choi 0001 |
VTC Fall | 2 |
| 2008 | Performance Analysis of IEEE802.11 Wireless Mesh NetworksabstractThe wireless mesh network is emerging as a promising technology in providing economical and scalable broadband Internet accesses to communities. The backbone of the wireless mesh network consists of mesh routers, which connect each other in an ad hoc manner via wireless links. The presence of backbone mesh routers and utilization of multiple channels and interfaces allow the wireless mesh network to have better capacity than that of the infrastructure-free ad hoc network formed by mesh clients directly. A special type of the mesh routers, referred to as gateway nodes, is capable of Internet connection, and other mesh routers and associated terminal clients have to access the Internet through the gateway nodes. In this paper, we present the analytically traceable stochastic models to characterize the average delay and throughput performance in wireless mesh networks. We model the forwarding mesh routers as an open queuing network. The analytical model takes into account the mesh router density, the random packet arrival process, the degree of locality of traffic and the collision avoidance mechanism of the IEEE802.11 DCF random access MAC. Our simulation results suggest that the analytical results are quite accurate, which can provide valuable insights in system performance and an effective guideline for the scalable design and optimization in wireless mesh networks. Hua Cai, Seung-Woo Seo |
ICC | 2 |
| 2008 | Performance measurement over Mobile WiMAX/IEEE 802.16e networkabstractWiBro(Wireless Broadband) is a Korean version of Mobile WiMAX/IEEE 802.16e system, which is designed for mobile broadband wireless access. Being a subset of IEEE 802.16e, WiBro employs orthogonal frequency division multiple access (OFDMA) and time division duplexing (TDD) schemes operating at 2.3 GHz bands. In mid 2006, the world first commercial Mobile WiMAX service, based on WiBro specification, started in Seoul, Korea. In this paper, we analyze the performance of commercial WiBro networks through traffic measurements. Many experiments are conducted in the various environments. We analyze the link capacity when the user datagram protocol (UDP) packets from multiple users fully utilize wireless links. We also analyze goodput performance with transmission control protocol (TCP) as well as round trip time performance. The measured performances are compared with those of HSDPA (High-Speed Downlink Packet Access) which is a competing system of Mobile WiMAX. We found that the RTTs of WiBro and HSDPA are very large compared with conventional data networks, e.g., Ethernet or Wireless LANs. Furthermore, it is shown that the user-perceived performance is limited by such long round trip times when TCP is utilized. It motivates us to improve round trip time performance of WiBro system. Finally, we analyze the VoIP performance and the performance when the user moves around the whole city. Dongmyoung Kim, Hua Cai, Minsoo Na, Sunghyun Choi 0001 |
WOWMOM | 2 |
| 2008 | A Roadside ITS Data Bus Prototype for Intelligent HighwaysabstractIntelligent transport systems (ITSs) usually include three principal elements: vehicle, driver, and road (or, more generally, the environment). The use of ITS data bus (IDB) has been proposed to build sharable standard interfaces for in-vehicle information systems. Despite broadband communication technologies, such as dedicated short-range communication (DSRC), which have been developed to provide high-quality roadside-vehicle communication services for intelligent highways, the existing IDB model has not paid enough attention to the demands of information exchanges between the roadside and onboard units. In this paper, a new model of the roadside IDB (RIDB) is proposed to improve the existing IDB architecture. The physical layer, data-link layer, and application layer of the new model are also discussed. A prototype system of the RIDB, which is based on the wireless 802.11b protocol, has been developed in a test site. The experimental results demonstrated that the RIDB is feasible to provide high-quality roadside-vehicle and roadside-roadside communication services. Other potential applications of the IDB, such as probe cars and intersection collision prevention, are also discussed. The RIDB proposed in this paper is potentially useful for the construction of an intelligent transport infrastructure. Hua Cai, Yingzi Lin |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2007 | An Epipolar Geometry-Based Fast Disparity Estimation Algorithm for Multiview Image and Video CodingabstractEffectively coding multiview visual content is an indispensable research topic because multiview image and video that provide greatly enhanced viewing experiences often contain huge amounts of data. Generally, conventional hybrid predictive-coding methodologies are adopted to address the compression by exploiting the temporal and interviewpoint redundancy existing in a multiview image or video sequences. However, their key yet time-consuming component, motion estimation (ME), is usually not efficient in interviewpoint prediction or disparity estimation (DE), because interviewpoint disparity is completely different from temporal motion existing in the conventional video. Targeting a generic fast DE framework for interviewpoint prediction, we propose a novel DE technique in this paper to accelerate the disparity search by employing epipolar geometry. Theoretical analysis, optimal disparity vector distribution histograms, and experimental results show that the proposed epipolar geometry-based DE can greatly reduce search region and effectively track large and irregular disparity, which is typical in convergent multiview camera setups. Compared with the existing state-of-the-art fast ME approaches, our proposed DE can obtain a similar coding efficiency while achieving a significant speedup for interviewpoint prediction and coding. Moreover, a robustness study shows that the proposed DE algorithm is insensitive to the epipolar geometry estimation noise. Hence, its wide application for multiview image and video coding is promising Jiangbo Lu, Hua Cai, Jianguang Lou, Jiang Li 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Multiview Image Coding Based on Geometric PredictionabstractMany existing multiview image/video coding techniques remove inter-viewpoint redundancy by applying disparity compensation in a conventional video coding framework, e.g., H.264/MPEG-4 AVC. However, conventional methodology works ineffectively as it ignores the special characteristics of inter-viewpoint disparity. In this paper, we propose a geometric prediction methodology for accurate disparity vector (DV) prediction, such that we can largely reduce the disparity compensation cost. Based on the new DV predictor, we design a basic framework that can be implemented in most existing multiview image/video coding schemes. We also use state-of-the-art H.264/MPEG-4 AVC as an example to illustrate how the proposed framework can be integrated with conventional video coding algorithms. Our experiments show proposed scheme can effectively tracks disparity and greatly improves coding performance. Compared with H.264/MPEG-4 AVC codec, our scheme outperforms maximally 1.5 dB when encoding some typical multiview image sequences. We also carry out an experiment to evaluate the robustness of our algorithm. The results indicate our method is robust and can be used in practical applications. Xing San, Hua Cai, Jianguang Lou, Jiang Li 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | An Effective Epipolar Geometry Assisted Motion Estimation Technique for Multi-View Image and Video CodingabstractTo efficiently encode data-intensive multi-view imaging content, conventional hybrid predictive coding methodologies choose to address the compression by exploiting temporal and inter-viewpoint redundancy. However, their key yet time-consuming component, motion estimation (ME), is usually not efficient in inter-viewpoint prediction because inter-viewpoint motion is quite different from temporal motion. In essence, inter-viewpoint correlation is subject to epipolar geometry, which provides constraints for multi-view image sequences. A fast inter-viewpoint ME technique is hence proposed in this paper to accelerate the encoding by employing epipolar geometry. Theoretical analysis and experimental results prove that the proposed ME algorithm can greatly reduce search region and effectively track large and irregular motion that is typical for convergent multi-view camera setups. As a result, compared with fast full search at large search size adopted in H.264, our proposed ME algorithm can obtain a similar coding efficiency while achieving a speedup ratio of 2.9. Jiangbo Lu, Hua Cai, Jian-Guang Lou, Jiang Li 0008 |
ICIP | 2 |
| 2006 | Color Image Coding by using Inter-Color CorrelationabstractInter-color correlation between the luminance component and chrominance components has been utilized for color image coding for years. However, the correlation has not been clearly analyzed. In this paper, we analyze the inter-color correlation and answer two questions related to color image coding: (1) what kind of inter-color correlation exists in color images after the discrete wavelet transform?; and, (2) how strong is it? This analysis helps us to find a most suitable inter-color context and eventually leads to a new embedded color image codec. By using the discovered inter-color context, significant performance improvement can be achieved when encoding chrominance components. Xing San, Hua Cai, Jiang Li 0008 |
ICIP | 2 |
| 2006 | Multicast of Real-Time Multi-View VideoabstractAs a recently emerging service, multi-view video provides a new viewing experience with high degree of freedom. However, due to the huge data amounts transferred, multi-view video's delivery remains a daunting challenge. In this paper, we propose a multi-view video-streaming system based on IP multicast. It can support a large number of users while still keeping a high degree of interactivity and low bandwidth consumption. Based on a careful user study, we have developed two schemes: one is for automatic delivery and the other for on-demand delivery. In automatic delivery, a server periodically multicasts special effect snapshots at a certain time interval. In on-demand delivery, the server delivers the snapshots based on distribution of users' requests. We conducted extensive experiments and user-experience studies to evaluate the proposed system's performance, and found that the system could provide satisfying multi-view video service for users on a large scale Li Zuo, Hua Cai, Jiang Li 0008 |
ICME | 3 |
| 2005 | Embedded image coding with context partitioning and quantizationabstractAs a key part of universal source coding, context quantization is very important for improving compression performance. However, in most existing methods, the quantizer is trained offline and is fixed due to the complexity of finding a good quantizer and the significant overhead of representing the quantizer. This paper proposes a novel online context quantization approach that achieves high coding efficiency with low quantizer overhead and computational complexity. It first partitions the context into groups according to the number of significant context events. A layer-based context quantization is then applied on these groups. The proposed method is applied for embedded wavelet image coding. Compared with the JPEG2000 coder, up to 0.6 dB improvements can be achieved on the standard 512 /spl times/ 512 test images. And more improvements are observed on images at lower resolutions. Hua Cai, Xing San, Jiang Li 0008 |
ICIP (2) | 1 |
| 2005 | Lossless image compression with tree coding of magnitude levelsabstractWith the rapid development of digital technology in consumer electronics, the demand to preserve raw image data for further editing or repeated compression is increasing. Traditional lossless image coders usually consist of computationally intensive modeling and entropy coding phases, therefore might not be suitable to mobile devices or scenarios with a strict real-time requirement. This paper presents a new image coding algorithm based on a simple architecture that is easy to model and encode the residual samples. In the proposed algorithm, each residual sample is separated into three parts: (1) a sign value, (2) a magnitude value, and (3) a magnitude level. A tree structure is then used to organize the magnitude levels. By simply coding the tree and the other two parts without any complicated modeling and entropy coding, good performance can be achieved with very low computational cost in the binary-uncoded mode. Moreover, with the aid of context-based arithmetic coding, the magnitude values are further compressed in the arithmetic-coded mode. This gives close performance to JPEG-LS and JPEG2000. Hua Cai, Jiang Li 0008 |
ICME | 1 |
| 2005 | A real-time interactive multi-view video systemabstractWith the rapid development of electronic and computing technology, multi-view video is attracting extensive interest recently due to its greatly enhanced viewing experience. In this paper, we present the system architecture for real-time capturing, processing, and interactive delivery of multi-view video. Unlike previous systems that mainly focus on multi-view video capturing, our system is designed to provide multi-view video service with high degree of interactivity in real time, which is still challenging in the current state of the technology. The proposed architecture tackles many practical problems in system calibration, object tracking, video compression, interactive delivery, etc. With the proposed system, users can interactively select their desired viewing directions and enjoy many exciting visual experiences, such as view switching, frozen moment and view sweeping, in real-time and with great freedom. Jian-Guang Lou, Hua Cai, Jiang Li 0008 |
ACM Multimedia | 2 |
| 2005 | A preliminary study on a fuzzy driving risk modelabstractOne potential problem of driver assistance systems is that they may lead to drivers' disengagement. Therefore, it is necessary to predict the danger ahead. Based on the discussion of the dangerous zone of a moving vehicle, driving risk level was proposed in this paper, and then a fuzzy logic based approach was presented to evaluate the risk level. The experiments of single-line motion demonstrated that the simulation result of driving risk level has the consistency of variance with the human driver's mental workload. The evaluation method of driving risk level has potential to be integrated into a DAS to sense the coming hazard like a human driver. Hua Cai |
SMC | 1 |
| 2004 | Error-resilient unequal protection of fine granularity scalable video bitstreamsabstractThis paper deals with the optimal packet loss protection issue for streaming the fine granularity scalable (FGS) video bitstreams over IP networks. Unlike many other existing protection schemes, we develop an error-resilient unequal protection (ER-UEP) method that adds redundant information optimally for loss protection and, at the same time, cancels completely the dependency among bitstream after loss recovery. In our ER-UEP method, the FGS enhancement-layer bitstream is first packetized into a group of independent data packets, while each packet can be truncated to represent the original video signal at any fidelity (i.e., scalability). Parity packets are then created with intrinsic UEP capabilities that can easily adapt to the current channel conditions. Unlike conventional UEP schemes that suffer from bitstream contamination due to the dependency among packets, our method guarantees the successful decoding of all received bits, thus leading to a better error resilience as well as higher robustness (under varying and/or unclean channel conditions). Hua Cai, Bing Zeng 0001, Guobin Shen, Shipeng Li 0001 |
ICC | 1 |
| 2004 | On packetization of JPEG2000 code-streams in wireless channelsabstractThe JPEG2000 standard adopts an error-resilient structure to organize its code-stream. Meanwhile, this robust structure also provides potential optimization space for packetization algorithms. In this paper, we study the optimal packetization issue for JPEG2000 code-streams over wireless channels. We first discuss the best strategy for packetization hits from an individual code-block into channel packets. Then, we extend it to bits from a group of code-blocks. Experimental results indicate that our scheme is outperforming the normal method by significant margins. Hua Cai |
ICIP | 1 |
| 2002 | Optimal rate allocation for macroblock-based progressive fine granularity scalable video codingabstractThis paper addresses the problem of optimal rate allocation for the macroblock-based progressive fine granularity scalable (PFGS) video coding. To solve this complicated problem, the error propagation pattern in the macroblock-based PFGS is first investigated. An effective drifting model is established subsequently for estimating the drifting for each enhancement bit stream segment encoded by the macroblock-based PFGS. The distortion reduction for the current frame and the estimated drifting suppression for the subsequent frames form the actual contribution of the enhancement layer bitstream. The equal-slope argument is then applied to select the best bit stream segments for the given bandwidth. Experiments show that our optimal rate allocation outperforms the uniform rate allocation by 0.3-1.4 dB. Hua Cai, Guobin Shen, Shipeng Li 0001, Bing Zeng 0001 |
ICIP (3) | 1 |
| 2002 | Error concealment for fine granularity scalable video transmissionabstractIn this paper we present an efficient error concealment (EC) method for the fine granularity scalable (FGS) video transmission. The proposed EC method exploits both the temporal and spatial correlations in an FGS encoded bitstream. In our scheme, the temporal redundancy is used to improve the quality of contaminated regions, and the intensity of the temporal correlation in contaminated regions is estimated by exploiting the spatial correlation in the surrounding high-quality regions. To maximally utilize the spatial correlation for estimation, we also propose two interleaving patterns that can avoid packetizing the neighboring regions into the same packet. Experiments show that our EC method achieves very good performance and is robust to different bandwidths and different sequences. Hua Cai, Guobin Shen, Feng Wu 0001, Shipeng Li 0001, Bing Zeng 0001 |
ICME (1) | 1 |