Wenhui Liao

dblp:26/6292 · DBLP profile ↗
← Back
21ranked-venue papers
12as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 10 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Information extraction and text analysis · 38% Vision and language · 33% Transfer learning and domain adaptation · 12%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%

Topics — the 20 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding · CVPR 2025
Natural language and speech › Information extraction and text analysis › document analysis
document information extraction
0.812024
PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking for End-to-end Document Pair Extraction · ACM Multimedia 2024
Natural language and speech › Information extraction and text analysis
entity linking
0.812024
PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking for End-to-end Document Pair Extraction · ACM Multimedia 2024
Machine learning › Transfer learning and domain adaptation › cross-domain learning
cross-domain recommendation
0.712023
A Deep Dual Adversarial Network for Cross-Domain Recommendation · IEEE Trans. Knowl. Data Eng. 2023
Recommender systems
cross-domain recommendation
0.712023
A Deep Dual Adversarial Network for Cross-Domain Recommendation · IEEE Trans. Knowl. Data Eng. 2023
Natural language and speech › Information extraction and text analysis › document analysis
document structure extraction
0.212024
PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking for End-to-end Document Pair Extraction · ACM Multimedia 2024
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › bayesian network
dynamic bayesian network
0.232007
Facial Action Unit Recognition by Exploiting Their Dynamic and Semantic Relationships · IEEE Trans. Pattern Anal. Mach. Intell. 2007
A Unified Probabilistic Framework for Facial Activity Modeling and Understanding · CVPR 2007
Inferring Facial Action Units with Causal Relations · CVPR (2) 2006
Computer vision › Face, body and person analysis
facial action unit recognition
0.122007
Facial Action Unit Recognition by Exploiting Their Dynamic and Semantic Relationships · IEEE Trans. Pattern Anal. Mach. Intell. 2007
Inferring Facial Action Units with Causal Relations · CVPR (2) 2006
Computer vision › Video understanding and tracking
object tracking
0.122007
Robust Object Tracking with a Case-Base Updating Strategy · IJCAI 2007
Robust Visual Tracking Using Case-Based Reasoning with Confidence · CVPR (1) 2006
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.112007
Facial Action Unit Recognition by Exploiting Their Dynamic and Semantic Relationships · IEEE Trans. Pattern Anal. Mach. Intell. 2007
Computer vision › Video understanding and tracking › object tracking
robust tracking
0.112007
Robust Object Tracking with a Case-Base Updating Strategy · IJCAI 2007
Computer vision › Face, body and person analysis › facial action unit analysis
AU relationship modeling
0.112006
Inferring Facial Action Units with Causal Relations · CVPR (2) 2006
Knowledge, reasoning and agents › Knowledge representation and reasoning
case-based reasoning
0.112006
Robust Visual Tracking Using Case-Based Reasoning with Confidence · CVPR (1) 2006
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
decision making under uncertainty
0.112006
Efficient Active Fusion for Decision-Making via VOI Approximation · AAAI 2006
Wearable and physiological sensing
emotion recognition
0.112006
Toward a decision-theoretic framework for affect recognition and user assistance · Int. J. Hum. Comput. Stud. 2006
Health and well-being technologies › health monitoring
stress recognition
0.112005
A Decision Theoretic Model for Stress Recognition and User Assistance · AAAI 2005
Usability and user experience research
user assistance
0.112005
A Decision Theoretic Model for Stress Recognition and User Assistance · AAAI 2005
Computer vision › Face, body and person analysis › facial expression analysis
facial expression recognition
0.012007
Facial Action Unit Recognition by Exploiting Their Dynamic and Semantic Relationships · IEEE Trans. Pattern Anal. Mach. Intell. 2007
Knowledge, reasoning and agents › Knowledge representation and reasoning › decision theory
decision-theoretic reasoning
0.012006
Toward a decision-theoretic framework for affect recognition and user assistance · Int. J. Hum. Comput. Stud. 2006
Computer vision › Face, body and person analysis
face tracking
0.012006
Robust Visual Tracking Using Case-Based Reasoning with Confidence · CVPR (1) 2006

Methods — techniques the papers use, named apart from their topics

orthogonal constraint · 1.3dual adversarial network · 1.3adversarial learning · 1.3multimodal integration · 0.9cot pre-training · 0.9chain-of-thought · 0.9transformer · 0.8LiLT · 0.8LayoutLMv3 · 0.8dynamic bayesian network · 0.2decision theory · 0.1
YearPublicationVenuePosition
2025 DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding
abstract
Text-rich document understanding (TDU) requires comprehensive analysis of documents containing substantial textual content and complex layouts. While Multimodal Large Language Models (MLLMs) have achieved fast progress in this domain, existing approaches either demand significant computational resources or struggle with effective multi-modal integration. In this paper, we introduce DocLayLLM, an efficient multi-modal extension of LLMs specifically designed for TDU. By lightly integrating visual patch tokens and 2D positional tokens into LLMs’ input and encoding the document content using the LLMs themselves, we fully take advantage of the document comprehension capability of LLMs and enhance their perception of OCR information. We have also deeply considered the role of chain-of-thought (CoT) and innovatively proposed the techniques of CoT Pre-training and CoT Annealing. Our DocLayLLM can achieve remarkable performances with lightweight training settings, showcasing its efficiency and effectiveness. Experimental results demonstrate that our DocLayLLM outperforms existing OCR-dependent methods and OCR-free competitors. Code and model are available at https://github.com/whlscut/DocLayLLM.
Wenhui Liao, Jiapeng Wang 0003, Chengyu Wang 0001, Jun Huang 0007
CVPR1
2024 ROISER: Towards Real World Semantic Entity Recognition from Visually-Rich Documents
Zening Lin, Wenhui Liao, Weicong Dai, Longfei Xiong
ICPR (31)3
2024 PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking for End-to-end Document Pair Extraction
abstract
Document pair extraction aims to identify key and value entities as well as their relationships from visually-rich documents. Most existing methods divide it into two separate tasks: semantic entity recognition (SER) and relation extraction (RE). However, simply concatenating SER and RE serially can lead to severe error propagation, and it fails to handle cases like multi-line entities in real scenarios. To address these issues, this paper introduces a novel framework, PEneo (Pair Extraction new decoder option), which performs document pair extraction in a unified pipeline, incorporating three concurrent sub-tasks: line extraction, line grouping, and entity linking. This approach alleviates the error accumulation problem and can handle the case of multi-line entities. Furthermore, to better evaluate the model's performance and to facilitate future research on pair extraction, we introduce RFUND, a re-annotated version of the commonly used FUNSD and XFUND datasets, to make them more accurate and cover realistic situations. Experiments on various benchmarks demonstrate PEneo's superiority over previous pipelines, boosting the performance by a large margin (e.g., 19.89%-22.91% F1 score on RFUND-EN) when combined with various backbones like LiLT and LayoutLMv3, showing its effectiveness and generality. Codes and the new annotations are available at https://github.com/ZeningLin/PEneo.
Zening Lin, Jiapeng Wang 0003, Wenhui Liao, Dayi Huang, Longfei Xiong
ACM Multimedia4
2023 A Deep Dual Adversarial Network for Cross-Domain Recommendation
abstract
Data sparsity is a common issue for most recommender systems and can severely degrade the usefulness of a system. One of the most successful solutions to this problem has been cross-domain recommender systems. These frameworks supplement the sparse data of the target domain with knowledge transferred from a source domain rich with data that is in some way related. However, there are three challenges that, if overcome, could significantly improve the quality and accuracy of cross-domain recommendation: 1) ensuring latent feature spaces of the users and items are both maximally matched; 2) taking consideration of user-item relationship and their interaction in modelling user preference; 3) enabling a two-way cross-domain recommendation that both the source and the target domains benefit from a knowledge exchange. Hence, in this paper, we propose a novel deep neural network called Dual Adversarial network for Cross-Domain Recommendation (DA-CDR). By training the shared encoders with a domain discriminator via dual adversarial learning, the latent feature spaces for both the users and items are maximally matched between the source and target domains. The domain-specific encoders are applied with an orthogonal constraint to ensure that any domain-specific features are properly extracted and work as supplement to the shared features. Allowing the two domains to collaboratively benefit from each other results in better recommendations for both domains. Extensive experiments with real-world datasets on six tasks demonstrate that DA-CDR significantly outperforms seven state-of-the-art baselines in terms of recommendation accuracy.
Qian Zhang 0023, Wenhui Liao, Guangquan Zhang 0001, Bo Yuan 0003, Jie Lu 0001
IEEE Trans. Knowl. Data Eng.2
2023 Heterogeneous Multidomain Recommender System Through Adversarial Learning
abstract
To solve the user data sparsity problem, which is the main issue in generating user preference prediction, cross-domain recommender systems transfer knowledge from one source domain with dense data to assist recommendation tasks in the target domain with sparse data. However, data are usually sparsely scattered in multiple possible source domains, and in each domain (source/target) the data may be heterogeneous, thus it is difficult for existing cross-domain recommender systems to find one source domain with dense data from multiple domains. In this way, they fail to deal with data sparsity problems in the target domain and cannot provide an accurate recommendation. In this article, we propose a novel multidomain recommender system (called HMRec) to deal with two challenging issues: 1) how to exploit valuable information from multiple source domains when no single source domain is sufficient and 2) how to ensure positive transfer from heterogeneous data in source domains with different feature spaces. In HMRec, domain-shared and domain-specific features are extracted to enable the knowledge transfer between multiple heterogeneous source and target domains. To ensure positive transfer, the domain-shared subspaces from multiple domains are maximally matched by a multiclass domain discriminator in an adversarial learning process. The recommendation in the target domain is completed by a matrix factorization module with aligned latent features from both the user and the item side. Extensive experiments on four cross-domain recommendation tasks with real-world datasets demonstrate that HMRec can effectively transfer knowledge from multiple heterogeneous domains collaboratively to increase the rating prediction accuracy in the target domain and significantly outperforms six state-of-the-art non-transfer or cross-domain baselines.
Wenhui Liao, Qian Zhang 0023, Bo Yuan 0003, Guangquan Zhang 0001, Jie Lu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2013 Stock Prediction Using Event-Based Sentiment Analysis
abstract
We propose a novel approach to label social media text using significant stock market events (big losses or gains). Since stock events are easily quantifiable using returns from indices or individual stocks, they provide meaningful and automated labels. We extract significant stock movements and collect appropriate pre, post and contemporaneous text from social media sources (for example, tweets from twitter). Subsequently, we assign the respective label (positive or negative) for each tweet. We train a model on this collected set and make predictions for labels of future tweets. We aggregate the net sentiment per each day (amongst other metrics) and show that it holds significant predictive power for subsequent stock market movement. We create successful trading strategies based on this system and find significant returns over other baseline methods.
Masoud Makrehchi, Sameena Shah, Wenhui Liao
Web Intelligence3
2009 Feature engineering on event-centric surrogate documents to improve search results
abstract
We investigate the task of re-ranking search results based on query log information. Prior work has considered this problem as either the task of learning document rankings of using features based on user behavior, or as the task of enhancing documents and queries using log data. Our contribution combines both. We distill log information into event-centric surrogate documents (ESDs), and extract features from these ESDs to be used in a learned ranking function. Our experiments on a legal corpus demonstrate that features engineered on surrogate documents lead to improved rankings, in particular when the original ranking is of poor quality.
Wenhui Liao, Isabelle Moulinier
CIKM1
2009 Learning Bayesian network parameters under incomplete data with domain knowledge
Wenhui Liao
Pattern Recognit.1
2009 Approximate Nonmyopic Sensor Selection via Submodularity and Partitioning
abstract
As sensors become more complex and prevalent, they present their own issues of cost effectiveness and timeliness. It becomes increasingly important to select sensor sets that provide the most information at the least cost and in the most timely and efficient manner. Two typical sensor selection problems appear in a wide range of applications. The first type involves selecting a sensor set that provides the maximum information gain within a budget limit. The other type involves selecting a sensor set that optimizes the tradeoff between information gain and cost. Unfortunately, both require extensive computations due to the exponential search space of sensor subsets. This paper proposes efficient sensor selection algorithms for solving both of these sensor selection problems. The relationships between the sensors and the hypotheses that the sensors aim to assess are modeled with Bayesian networks, and the information gain (benefit) of the sensors with respect to the hypotheses is evaluated by mutual information. We first prove that mutual information is a submodular function in a relaxed condition, which provides theoretical support for the proposed algorithms. For the budget-limit case, we introduce a greedy algorithm that has a constant factor of$(1 - 1/e)$guarantee to the optimal performance. A partitioning procedure is proposed to improve the computational efficiency of the algorithms by efficiently computing mutual information as well as reducing the search space. For the optimal-tradeoff case, a submodular–supermodular procedure is exploited in the proposed algorithm to choose the sensor set that achieves the optimal tradeoff between the benefit and cost in a polynomial-time complexity.
Wenhui Liao, William A. Wallace
IEEE Trans. Syst. Man Cybern. Part A1
2008 Exploiting qualitative domain knowledge for learning Bayesian network parameters with incomplete data
abstract
When a large amount of data are missing, or when multiple hidden nodes exist, learning parameters in Bayesian networks (BNs) becomes extremely difficult. This paper presents a learning algorithm to incorporate qualitative domain knowledge to regularize the otherwise ill-posed problem, limit the search space, and avoid local optima. Specifically, the problem is formulated as a constrained optimization problem, where an objective function is defined as a combination of the likelihood function and penalty functions constructed from the qualitative domain knowledge. Then, a gradient-descent procedure is systematically integrated with the E-step and M-step of the EM algorithm, to estimate the parameters iteratively until it converges. The experiments show our algorithm improves the accuracy of the learned BN parameters significantly over the conventional EM algorithm.
Wenhui Liao
ICPR1
2008 Efficient non-myopic value-of-information computation for influence diagrams
Wenhui Liao
Int. J. Approx. Reason.1
2007 A Unified Probabilistic Framework for Facial Activity Modeling and Understanding
abstract
Facial activities are the most natural and powerful means of human communication. Spontaneous facial activity is characterized by rigid head movements, non-rigid facial muscular movements, and their interactions. Current research in facial activity analysis is limited to recognizing rigid or non-rigid motion separately, often ignoring their interactions. Furthermore, although some of them analyze the temporal properties of facial features during facial feature extraction, they often recognize the facial activity statically, ignoring the dynamics of the facial activity. In this paper, we propose to explicitly exploit the prior knowledge about facial activities and systematically combine the prior knowledge with image measurements to achieve an accurate, robust, and consistent facial activity understanding. Specifically, we propose a unified probabilistic framework based on the dynamic Bayesian network (DBN) to simultaneously and coherently represent the rigid and non-rigid facial motions, their interactions, and their image observations, as well as to capture the temporal evolution of the facial activities. Robust computer vision methods are employed to obtain measurements of both rigid and non-rigid facial motions. Finally, facial activity recognition is accomplished through a probabilistic inference by systemically integrating the visual measurements with the facial activity model.
Wenhui Liao, Zheng Xue
CVPR2
2007 Robust Object Tracking with a Case-Base Updating Strategy
Wenhui Liao
IJCAI1
2007 Facial Action Unit Recognition by Exploiting Their Dynamic and Semantic Relationships
abstract
A system that could automatically analyze the facial actions in real time has applications in a wide range of different fields. However, developing such a system is always challenging due to the richness, ambiguity, and the dynamic nature of facial actions. Although a number of research groups attempt to recognize facial action units (AUs) by either improving facial feature extraction techniques, or the AU classification techniques, these methods often recognize AUs or certain AU combinations individually and statically, ignoring the semantic relationships among AUs and the dynamics of AUs. Hence, these approaches cannot always recognize AUs reliably, robustly, and consistently. In this paper, we propose a novel approach that systematically accounts for the relationships among AUs and their temporal evolutions for AU recognition. Specifically, we use a dynamic Bayesian network (DBN) to model the relationships among different AUs. The DBN provides a coherent and unified hierarchical probabilistic framework to represent probabilistic relationships among various AUs and to account for the temporal changes in facial action development. Within our system, robust computer vision techniques are used to obtain AU measurements. And such AU measurements are then applied as evidence to the DBN for inferring various AUs. The experiments show that the integration of AU relationships and AU dynamics with AU measurements yields significant improvement of AU recognition, especially for spontaneous facial expressions and under more realistic environment including illumination variation, face pose variation, and occlusion.
Wenhui Liao
IEEE Trans. Pattern Anal. Mach. Intell.2
2006 Efficient Active Fusion for Decision-Making via VOI Approximation
Wenhui Liao
AAAI1
2006 Inferring Facial Action Units with Causal Relations
abstract
A system that could automatically analyze the facial actions in real time have applications in a number of different fields. However, developing such a system is always a challenging task due to the richness, ambiguity, and dynamic nature of facial actions. Although a number of research groups attempt to recognize action units (AUs) by either improving facial feature extraction techniques, or the AU classification techniques, these methods often recognize AUs individually and statically, therefore ignoring the semantic relationships among AUs and the dynamics of AUs. Hence, these approaches cannot always recognize AUs reliably, robustly, and consistently. In this paper, we propose a novel approach for AUs classification, that systematically accounts for relationships among AUs and their temporal evolution. Specifically, we use a dynamic Bayesian network (DBN) to model the relationships among different AUs. The DBN provides a coherent and unified hierarchical probabilistic framework to represent probabilistic relationships among different AUs and account for the temporal changes in facial action development. Under our system, robust computer vision techniques are used to get AU measurements. And such AU measurements are then applied as evidence into the DBN for inferencing various AUs. The experiments show the integration of AU relationships and AU dynamics with AU image measurements yields significant improvements in AU recognition.
Wenhui Liao
CVPR (2)2
2006 Robust Visual Tracking Using Case-Based Reasoning with Confidence
abstract
The paper describes a simple but robust framework for visual object tracking in a video sequence. Compared with the existing tracking techniques, our proposed tracking technique has two significant contributions. First, a Case- Based Reasoning (CBR) paradigm is introduced to track the non-rigid object robustly under significant appearance changes without drifting away. Second, it can provide an accurate confidence measurement for each tracked object so that the tracking failures can be identified successfully. Specifically, under this framework, the appearance changes of the object being tracked can be adapted dynamically during tracking via an adaption mechanism of CBR. Hence, an accurate 2D tracking model can be maintained online for each image frame during tracking. Therefore, the proposed tracking technique possesses a self-recovery capability so that the object can be tracked robustly under significant appearance changes without error accumulation. Application was focused on the development of a real-time face tracking system. Via the proposed framework, the built real-time face tracker can track the human face robustly at 26 frames per second under various face orientations, significant facial expression and external illumination changes.
Wenhui Liao
CVPR (1)2
2006 Toward a decision-theoretic framework for affect recognition and user assistance
Wenhui Liao, Wayne D. Gray
Int. J. Hum. Comput. Stud.1
2005 A Decision Theoretic Model for Stress Recognition and User Assistance
Wenhui Liao
AAAI1
2004 A Factor Tree Inference Algorithm for Bayesian Networks and Its Application
abstract
In a Bayesian network, a probabilistic inference is the procedure of computing the posterior probability of query variables given a collection of evidences. In This work, we propose an algorithm that efficiently carries out the inferences whose query variables and evidence variables are restricted to a subset of the set of the variables in a BN. The algorithm successfully combines the advantages of two popular inference algorithms - variable elimination and clique tree propagation. We empirically demonstrate its computational efficiency in an affective computing domain.
Wenhui Liao
ICTAI1
2002 Scene change detection by audio and video clues
abstract
Automatic video scene change detection is a challenging task. Using audio or visual information alone often cannot provide a satisfactory solution. However, how to combine audio and visual information efficiently still remains a difficult issue since there are various cases in their relationship due to the versatility of videos. We present an effective scene change detection method that adopts the joint evaluation of the audio and visual features. First, video information is used to find the shot boundaries. Second, the audio features for each video shot can be extracted. Lastly, an audio-video combination schema is proposed to detect the video scene boundaries.
Shu-Ching Chen, Mei-Ling Shyu, Wenhui Liao, Chengcui Zhang
ICME (2)3