Zhidong Li

dblp:10/6710 · DBLP profile ↗
← Back
53ranked-venue papers
7as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 15 · 7 since 2021Human-computer interaction and ubiquitous computing · 4Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RcAE: Recursive Reconstruction Framework for Unsupervised Industrial Anomaly Detection
abstract
Unsupervised industrial anomaly detection requires accurately identifying defects without labeled data. Traditional autoencoder-based methods often struggle with incomplete anomaly suppression and loss of fine details, as their single-pass decoding fails to effectively handle anomalies with varying severity and scale. We propose a recursive architecture for autoencoder (RcAE), which performs reconstruction iteratively to progressively suppress anomalies while refining normal structures. Unlike traditional single-pass models, this recursive design naturally produces a sequence of reconstructions, progressively exposing suppressed abnormal patterns. To leverage this reconstruction dynamics, we introduce a Cross Recursion Detection (CRD) module that tracks inconsistencies across recursion steps, enhancing detection of both subtle and large-scale anomalies. Additionally, we incorporate a Detail Preservation Network (DPN) to recover high-frequency textures typically lost during reconstruction. Extensive experiments demonstrate that our method significantly outperforms existing non-diffusion methods, and achieves performance on par with recent diffusion models with only 10% of their parameters and offering substantially faster inference. These results highlight the practicality and efficiency of our approach for real-world applications.
Rongcheng Wu, Shiying Zhang, Zhidong Li, Hui Li 0005, Jianlong Zhou, Jiangtao Cui, Fang Chen 0001, Pingyang Sun, Qiyu Liao
AAAI5
2026 Can GNNs Learn Link Heuristics? a Concise Review and Evaluation of Link Prediction Methods
abstract
This paper explores the ability of Graph Neural Networks (GNNs) in learning various forms of information for link prediction, alongside a brief review of existing link prediction methods. Our analysis reveals that GNNs cannot effectively learn structural information related to the number of common neighbors between two nodes, primarily due to the nature of set-based pooling of the neighborhood aggregation scheme. Also, our extensive experiments indicate that trainable node embeddings can improve the performance of GNN-based link prediction models. Importantly, we observe that the denser the graph, the greater such the improvement. We attribute this to the characteristics of node embeddings, where the link state of each link sample could be encoded into the embeddings of nodes that are involved in the neighborhood aggregation of the two nodes in that link sample. In denser graphs, every node could have more opportunities to attend the neighborhood aggregation of other nodes and encode states of more link samples to its embedding, thus learning better node embeddings for link prediction. Lastly, we demonstrate that the insights gained from our research carry important implications in identifying the limitations of existing link prediction methods, which could guide the future development of more robust algorithms.
Shuming Liang, Zhidong Li, Bin Liang 0003, Yang Wang 0002, Fang Chen 0001
IEEE Trans. Big Data3
2026 Deep Learning-Based Instance Segmentation: A Comprehensive Review of Algorithms, Challenges, and Future Directions
Jiacheng Lou, Sergei Shavetov, Xuecheng Wen, Zhidong Li, Chunhong Yuan
Vis. Comput.4
2025 Navigating Towards Fairness with Data Selection
abstract
Machine learning algorithms often struggle to eliminate inherent data biases, particularly those arising from unreliable labels, which poses a significant challenge in ensuring fairness. Existing fairness techniques that address label bias typically involve modifying models and intervening in the training process, but these lack flexibility for large-scale datasets. To address this limitation, we introduce a data selection method designed to efficiently and flexibly mitigate label bias, tailored to more practical needs. Our approach utilizes a zero-shot predictor as a proxy model that simulates training on a clean holdout set. This strategy, supported by peer predictions, ensures the fairness of the proxy model and eliminates the need for an additional holdout set, which is a common requirement in previous methods. Without altering the classifier's architecture, our modality-agnostic method effectively selects appropriate training data and has proven efficient and effective in handling label bias and improving fairness across diverse datasets in experimental evaluations.
Yixuan Zhang 0006, Zhidong Li, Yang Wang 0002, Fang Chen 0001, Xuhui Fan 0001, Feng Zhou 0011
AAAI2
2025 Spatio-Temporal Residual Masked Autoencoder for Urban Rent Estimation
Chenya Huang, Bin Liang 0003, Zhidong Li, Justin Wang, Fang Chen 0001
CIKM3
2025 Linking Cryptoasset Attribution Tags to Knowledge Graph Entities: An LLM-Based Approach
Régnier Avice, Bernhard Haslhofer, Zhidong Li, Jianlong Zhou
FC (2)3
2025 Multimodal Machine Learning for Real Estate Appraisal: A Comprehensive Survey
Chenya Huang, Bin Liang 0003, Zhidong Li, Fang Chen 0001
PAKDD (4)3
2025 Toward fair medical advice: Addressing and mitigating bias in large language model-based healthcare applications
abstract
Large Language Models (LLMs) are increasingly deployed in web-based medical advice applications, offering scalable and accessible healthcare solutions. However, their outputs often reflect demographic biases, raising concerns about fairness and equity for vulnerable populations. In this work, we propose FairMed, a framework designed to mitigate biases in LLM-generated medical advice through fine-tuning and prompt engineering strategies. We evaluate FairMed using language-based and content-level metrics across demographic groups on publicly available (MedQA), synthetic (Synthea), and private (CBHS) datasets. Experimental results demonstrate consistent improvements over Llama3 - Med42, as well as over the zero-shot prompting baseline. For instance, in sentiment analysis for gender groups using MedQA, FairMed with Descriptive Prompting reduces the Statistical Parity Difference (SPD) from 0.0902 to 0.0658, improves the Disparate Impact Ratio from 1.1916 to 1.1566, and decreases the Kullback-Leibler Divergence from 0.0045 to 0.0024. Similarly, in directive language evaluation for gender groups using Synthea, SPD improves from 0.1056 to nearly zero, achieving near-perfect parity. On the CBHS dataset, FairMed with Descriptive Prompting increases Diagnostic Recommendation Divergence (DRD) for race groups from 0.9530 to 0.9848, indicating improved group-specific tailoring, while reducing the Action Disparity Index (ADI) from 0.0857 to 0.0469 and Referral Frequency Parity (RFP) from 0.0791 to 0.0511, reflecting enhanced fairness. These findings highlight FairMed's effectiveness in addressing demographic disparities and promoting equitable healthcare guidance through web technologies. This framework contributes to building trustworthy and inclusive systems for delivering medical advice by ensuring fairness in sensitive applications.
Haohui Lu, Zhidong Li, Man Lung Yiu, Yu Gao 0025, Shahadat Uddin
Artif. Intell. Medicine3
2024 Expert-Guided Model Cultivation: CoTeaching to Resolve Abstruseness and Enhance Learning Performance
Feng Zhou 0011, Zhidong Li, Yang Wang 0002, Donglian Qi, Shuming Li
ADMA (2)3
2024 TransFeat-TPP: An Interpretable Deep Covariate Temporal Point Processes
abstract
The classical temporal point process (TPP) constructs an intensity function by taking the occurrence times into account. Nevertheless, occurrence time may not be the only relevant factor, other contextual data, termed covariates, may also impact the event evolution. Incorporating such covariates into the model is beneficial, while distinguishing their relevance to the event dynamics is of great practical significance. In this work, we propose a Transformer-based covariate temporal point process (TransFeat-TPP) model to improve the interpretability of deep covariate-TPPs while maintaining powerful expressiveness. TransFeat-TPP can effectively model complex relationships between events and covariates, and provide enhanced interpretability by discerning the importance of various covariates. Experimental results on synthetic and real datasets demonstrate improved prediction accuracy and consistently interpretable feature importance when compared to existing deep covariate-TPPs. Our code is available at https://github.com/waystogetthere/TransFeat.git.
Zizhuo Meng, Boyu Li 0003, Xuhui Fan 0001, Zhidong Li, Yang Wang 0002, Fang Chen 0001, Feng Zhou 0011
ECAI4
2024 VAGNN: Advancing the Generalization of Graph Neural Networks
Shuming Liang, Bin Liang 0003, Zhidong Li, Yang Wang 0002, Fang Chen 0001
ICONIP (1)4
2024 Interpretable Transformer Hawkes Processes: Unveiling Complex Interactions in Social Networks
abstract
Social networks represent complex ecosystems where the interactions between users or groups play a pivotal role in information dissemination, opinion formation, and social interactions.Effectively harnessing event sequence data within social networks to unearth interactions among users or groups has persistently posed a challenging frontier within the realm of point processes.Current deep point process models face inherent limitations within the context of social networks, constraining both their interpretability and expressive power.These models encounter challenges in capturing interactions among users or groups and often rely on parameterized extrapolation methods when modeling intensity over non-event intervals, limiting their capacity to capture complex intensity patterns beyond observed events.To address these challenges, this study proposes modifications to Transformer Hawkes processes (THP), leading to the development of interpretable Transformer Hawkes processes (ITHP).ITHP inherits the strengths of THP while aligning with statistical nonlinear Hawkes processes, thereby enhancing its interpretability and providing valuable insights into interactions between users or groups.Additionally, ITHP enhances the flexibility of the intensity function over non-event intervals, making it better suited to capture complex event propagation patterns in social networks.Experimental results, both on synthetic and real data, demonstrate the effectiveness of ITHP in overcoming the identified limitations.Moreover, they highlight ITHP's applicability in the context of exploring the complex impact
Zizhuo Meng, Ke Wan 0002, Yadong Huang, Zhidong Li, Yang Wang 0002, Feng Zhou 0011
KDD4
2024 A model-driven dual-derivation framework for quantitative fault detection in satellite power system
Pengming Wang 0002, Liansheng Liu, Zhidong Li, Datong Liu
Adv. Eng. Informatics4
2024 Few-Shot Stereo Matching with High Domain Adaptability Based on Adaptive Recursive Network
Rongcheng Wu, Zhidong Li, Jianlong Zhou, Fang Chen 0001, Changming Sun
Int. J. Comput. Vis.3
2023 Fair Representation Learning with Unreliable Labels
abstract
In learning with fairness, for every instance, its label can be randomly flipped to another class due to the practitioner’s prejudice, namely, label bias. The existing well-studied fair representation learning methods focus on removing the dependency between the sensitive factors and the input data, but do not address how the representations retain useful information when the labels are unreliable. In fact, we find that the learned representations become random or degenerated when the instance is contaminated by label bias. To alleviate this issue, we investigate the problem of learning fair representations that are independent of the sensitive factors while retaining the task-relevant information given only access to unreliable labels. Our model disentangles the dependency between fair representations and sensitive factors in the latent space. To remove the reliance between the labels and sensitive factors, we incorporate an additional penalty based on mutual information. The learned purged fair representations can then be used in any downstream processing. We demonstrate the superiority of our method over previous works through multiple experiments on both synthetic and real-world datasets.
Yixuan Zhang 0006, Feng Zhou 0011, Zhidong Li, Yang Wang 0002, Fang Chen 0001
AISTATS3
2023 Heterogeneous domain adaptation by semantic distribution alignment network
Weihua Jin, Pengming Wang 0002, Bo Sun 0013, Zhidong Li
Appl. Intell.5
2021 Bias-tolerant Fair Classification
abstract
The label bias and selection bias are acknowledged as two reasons in data that will hinder the fairness of machine-learning outcomes. The label bias occurs when the labeling decision is disturbed by sensitive features, while the selection bias occurs when subjective bias exists during the data sampling. Even worse, models trained on such data can inherit or even intensify the discrimination. Most algorithmic fairness approaches perform an empirical risk minimization with predefined fairness constraints, which tends to trade-off accuracy for fairness. However, such methods would achieve the desired fairness level with the sacrifice of the benefits (receive positive outcomes) for individuals affected by the bias. Therefore, we propose a \textbf{B}ias-Tolerant \textbf{FA}ir \textbf{R}egularized \textbf{L}oss (B-FARL), which tries to regain the benefits using data affected by label bias and selection bias. B-FARL takes the biased data as input, calls a model that approximates the one trained with fair but latent data, and thus prevents discrimination without constraints required. In addition, we show the effective components by decomposing B-FARL, and we utilize the meta-learning framework for the B-FARL optimization. The experimental results on real-world datasets show that our method is empirically effective in improving fairness towards the direction of true but latent labels.
Yixuan Zhang 0006, Feng Zhou 0011, Zhidong Li, Yang Wang 0002, Fang Chen 0001
ACML3
2021 Failure Prediction for Large-scale Water Pipe Networks Using GNN and Temporal Failure Series
abstract
Pipe failure prediction in the water industry aims to prioritize the pipes that are at high risk of failure for proactive maintenance. However, existing statistical or machine learning models that rely on historical failures and asset attributes can hardly leverage the structure information of pipe networks. In this work, we develop a failure prediction framework for pipe networks by jointly considering the pipes' features, the network structure, the geographical neighboring effect, and the temporal failure series. We apply a multi-hop Graph Neural Network (GNN) to failure prediction. We propose a method of constructing a geographical graph structure depending on not only the physical connections but also geographical distances between pipes. To differentiate the pipes with diverse properties, we employ an attention mechanism in the neighborhood aggregation process of each GNN layer. Also, residual connections and layer-wise aggregation are used to avoid the over-smoothing issue in deep GNNs. The historical failures exhibit a strong temporal pattern. Inspired by point process, we develop a module to learn the pipes' evolutionary effect and the time-decayed excitement of historical failures on the current state of the pipe. The proposed framework is evaluated on two real-world large-scale pipe networks. It outperforms the existing statistical, machine learning, and state-of-the-art GNN baselines. Our framework provides the water utility with core data-driven support for proactive maintenance including regular pipe inspection, pipe renewal planning, and sensor system deployment. It can be extended to other infrastructure networks in the future.
Shuming Liang, Zhidong Li, Bin Liang 0003, Yang Wang 0002, Fang Chen 0001
CIKM2
2021 A Multi-task Kernel Learning Algorithm for Survival Analysis
Zizhuo Meng, Jie Xu 0008, Zhidong Li, Yang Wang 0002, Fang Chen 0001, Zhiyong Wang 0001
PAKDD (3)3
2020 A Data Driven Approach for Leak Detection with Smart Sensors
abstract
Preventing water pipe leaks and breaks has high priority for water utilities. It is a critical task for the utility to reduce water loss through leaks and breaks detection in water mains. The failure prediction and data analytics research have been conducted for an Australian water utility over the last few years to enhance the prediction of leaks and breaks detection in water mains. Intelligent sensing at sensitive locations with current research aids in prioritising investigation and prevention of potential breaks and leaks in water mains. The purpose of this work is to integrate the predictive analytics and intelligent sensing applications to identify high risk mains prior to failures. Predictive analytics and minimum night flow (MNF) analysis have been utilised to prioritise risky zones over the whole water network, and then risky pipes are identified to optimise sensors deployment. The sensing data is being collected for analysis and validation, and a machine learning model is being built based on the analysis results. This work is currently under progress and the planned outcomes will help the utility reduce water loss, improve leak detection, and enhance customer satisfaction by automating the process of leak detection using a data driven approach with smart sensors.
Bin Liang 0003, Sunny Verma, Jie Xu 0008, Shuming Liang, Zhidong Li, Yang Wang 0002, Fang Chen 0001
ICARCV5
2020 Simultaneous Customer Segmentation and Behavior Discovery
Ling Luo 0002, Zhidong Li, Yang Wang 0002, Fang Chen 0001
ICONIP (4)3
2020 Long-Term Pipeline Failure Prediction Using Nonparametric Survival Analysis
Dilusha Weeraddana, Sudaraka Mallawaarachchi, Tharindu Warnakula, Zhidong Li, Yang Wang 0002
ECML/PKDD (4)4
2019 Physiological Indicators for User Trust in Machine Learning with Influence Enhanced Fact-Checking
Jianlong Zhou, Huaiwen Hu, Zhidong Li, Kun Yu 0001, Fang Chen 0001
CD-MAKE3
2019 Predicting Water Quality for the Woronora Delivery Network with Sparse Samples
abstract
Monitoring drinking water quality in the entire delivery network, mainly indicated by total chlorine (TC), is a critical component of overall water supply management. However, it is extremely difficult to collect sufficient TC data from the network at customer sites, which makes it sparse for comprehensive modelling. This paper details an approach that provides TC prediction within the entire Woronora delivery network in Sydney in the next 24 hours. First, the hydraulic system is employed to capture the topology of the delivery network, so that the water travel time can be estimated using predicted water demand. The travel time links the upstream (reservoir) data to the downstream (resident) data. Then, a two-step strategy is proposed as a semi-parametric method to determine the crucial factors and build Bayesian model for TC decay to predict TC with the travel time. Lastly, the uncertainties of both data and the model are analysed to define the boundaries of prediction for better decision making. Several operational stages are involved when the approach is being deployed, including prediction interpretation, interactive tool development for water quality mapping and visualisation, and proactive optimisation. This has established a successful initiative to improve the overall water supply management for the entire Woronora delivery network.
Bin Liang 0003, Dammika Vitanage, Corinna Doolan, Zhidong Li, Ronnie Taib, George Mathews, Yang Wang 0002, Shiyang Lu, Fang Chen 0001, Tin Hua, Andrew Peters
ICDM4
2019 Exploring Latent Structure Similarity for Bayesian Nonparameteric Model with Mixture of NHPP Sequence
Yongzhe Chang, Zhidong Li, Ling Luo 0002, Simon Luo, Arcot Sowmya, Yang Wang 0002, Fang Chen 0001
ICONIP (2)2
2019 Recovering DTW Distance Between Noise Superposed NHPP
Yongzhe Chang, Zhidong Li, Bang Zhang, Ling Luo 0002, Arcot Sowmya, Yang Wang 0002, Fang Chen 0001
PAKDD (2)2
2019 Multitask Learning for Sparse Failure Prediction
Simon Luo, Victor W. Chu, Zhidong Li, Yang Wang 0002, Jianlong Zhou, Fang Chen 0001, Raymond K. Wong 0001
PAKDD (1)3
2019 Hawkes Process with Stochastic Triggering Kernel
Feng Zhou 0011, Yixuan Zhang 0006, Zhidong Li, Xuhui Fan 0001, Yang Wang 0002, Arcot Sowmya, Fang Chen 0001
PAKDD (1)3
2019 Multi-objective optimization of energy consumption in crude oil pipeline transportation system operation based on exergy loss analysis
Yang Liu 0079, Qinglin Cheng, Yifan Gan, Zhidong Li
Neurocomputing5
2018 Long-Term RNN: Predicting Hazard Function for Proactive Maintenance of Water Mains
abstract
Failure event prediction is becoming increasingly important in wide applications, such as the planning of proactive maintenance, the active investment management, and disease surveillance. To address the issue, the hazard function in survival analysis has been employed to describe the pattern of failures. Different from traditional survival analysis, this paper discovers how to apply recurrent neural network (RNN) to the long-term hazard function prediction. The proposed Long-Term RNN (LT-RNN) is able to leverage the precedent information shared by other entities, leading to more reliable long-term predictions. Specifically, our method allows a black-box treatment for modelling the hazard function which is often a pre-defined parametric form in typical survival analysis. The key idea of our approach is to model the hazard function as a nonparameteric function of the history. The same precedent information from other entities is embedded to a stitched vector for LT-RNN to automatically learn a representation of the long-term hazard function. We apply our model to the proactive maintenance problem using a large dataset from a water utility in Australia.
Bin Liang 0003, Zhidong Li, Yang Wang 0002, Fang Chen 0001
CIKM2
2018 A Refined MISD Algorithm Based on Gaussian Process Regression
Feng Zhou 0011, Zhidong Li, Xuhui Fan 0001, Yang Wang 0002, Arcot Sowmya, Fang Chen 0001
PAKDD (2)2
2018 End-User Development for Interactive Data Analytics: Uncertainty, Correlation and User Confidence
abstract
This paper investigates End-User Development (EUD) for interactive data-analytic interfaces-building upon the ideas of making machine learning transparent. The research is carried out in a business operation environment (water pipe failure prediction in our case) motivated to integrate advanced analytics into decision-making processes of an urban Internet of Things (IoT) concept. We explore effects of revealing uncertainty and correlation on user confidence in a data-driven decision making scenario. It was found that user confidence varied significantly amongst various user groups when different machine learning models were displayed with/without supplementary information. Galvanic Skin Response (GSR) signals were analyzed and shown as reasonable indices for predicting user confidence levels. Supplementary data visualizations (of inherent uncertainty and correlation in data) contributed to explicability principles while GSR indexing added towards correctibility principles. We recommend transparent machine learning as the key to effective EUD for interactive data analytics.
Jianlong Zhou, Syed Arshad, Xiuying Wang 0001, Zhidong Li, David Dagan Feng, Fang Chen 0001
IEEE Trans. Affect. Comput.4
2016 Robust Bayesian non-parametric dictionary learning with heterogeneous Gaussian noise
Yi Wang 0041, Bin Li 0015, Yang Wang 0002, Fang Chen 0001, Bang Zhang, Zhidong Li
Comput. Vis. Image Underst.6
2015 Data Driven Water Pipe Failure Prediction: A Bayesian Nonparametric Approach
abstract
Water pipe failures can cause significant economic and social costs, hence have become the primary challenge to water utilities. In this paper, we propose a Bayesian nonparametric approach, namely the Dirichlet process mixture of hierarchical beta process model, for water pipe failure prediction. It can select high-risk pipes for physical condition assessment, thereby preventing disastrous failures proactively.
Bang Zhang, Yi Wang 0041, Zhidong Li, Bin Li 0015, Yang Wang 0002, Fang Chen 0001
CIKM4
2015 Measurable Decision Making with GSR and Pupillary Analysis for Intelligent User Interface
abstract
This article presents a framework of adaptive, measurable decision making for Multiple Attribute Decision Making (MADM) by varying decision factors in their types, numbers, and values. Under this framework, decision making is measured using physiological sensors such as Galvanic Skin Response (GSR) and eye-tracking while users are subjected to varying decision quality and difficulty levels. Following this quantifiable decision making, users are allowed to refine several decision factors in order to make decisions of high quality and with low difficulty levels. A case study of driving route selection is used to set up an experiment to test our hypotheses. In this study, GSR features exhibit the best performance in indexing decision quality. These results can be used to guide the design of intelligent user interfaces for decision-related applications in HCI that can adapt to user behavior and decision-making performance.
Jianlong Zhou, Jinjun Sun, Fang Chen 0001, Yang Wang 0002, Ronnie Taib, Ahmad Khawaji, Zhidong Li
ACM Trans. Comput. Hum. Interact.7
2014 Water pipe condition assessment: a hierarchical beta process approach for sparse incident data
Zhidong Li, Bang Zhang, Yang Wang 0002, Fang Chen 0001, Ronnie Taib, Vicky Whiffin, Yi Wang 0041
Mach. Learn.1
2013 The Effect of Stress on Cognitive Load Measurement
Dan Conway, Ian Dick, Zhidong Li, Yang Wang 0002, Fang Chen 0001
INTERACT (4)3
2013 Indexing cognitive workload based on pupillary response under luminance and emotional changes
abstract
Pupillary response is a popular physiological index of cognitive workload that can be used for design and evaluation of adaptive interface in various areas of human-computer interaction (HCI) research. However, in practice various confounding factors unrelated to workload, including changes of luminance condition and emotional arousal might degrade pupillary response based workload measures such as commonly used mean pupil diameter. This work investigates pupillary response as a cognitive workload measure under the influence of such confounding factors. Video-based eye tracker is used to record pupillary response during arithmetic tasks under luminance and emotional changes. Machine learning based feature selection and classification techniques are proposed to robustly index cognitive workload based on pupillary response even with the influence of noisy factors unrelated to workload.
Zhidong Li, Yang Wang 0002, Fang Chen 0001
IUI2
2013 A Bayesian non-parametric viewpoint to visual tracking
abstract
A novel bayesian non-parametric method for tracking is proposed in this paper. The foreground appearance distribution is modeled by unbounded mixtures controlled through a Bayesian non-parametric process. Two posterior inference strategies are provided: Gibbs sampling and sequential importance sampling. Both of these two sampling strategy benefits from the conjugate prior/posterior pairs by factorizing the joint posterior distributions. Once the mixture model is obtained/updated, the similarities/probablity of each observations assigned to this mixture model could be easily calculated. In model matching/verification, the Kullback-Leibler divergence and texture information is adopted for verification purpose. The robustness of our methods is demonstrated by the experiments.
Yi Wang 0041, Zhidong Li, Yang Wang 0002, Fang Chen 0001
WACV2
2013 Visual tracking by proto-objects
Zhidong Li, Yang Wang 0002, Fang Chen 0001, Yi Wang 0041
Pattern Recognit.1
2012 New Robust H ∞ Fuzzy Control for the Interconnected Bilinear Systems Subject to Actuator Saturation
Dongsheng Yang 0001, Zhidong Li
ISNN (2)3
2011 Saliency detection based on proto-objects and topic model
abstract
This paper proposes a novel computational framework for saliency detection, which integrates the saliency map computation and proto-objects detection. The proto-objects are detected based on the saliency map using latent topic model. The detected proto-objects are then utilized to improve the saliency map computation. Extensive experiments are performed on two publicly available datasets. The experimental results show that the proposed framework outperforms the state-of-art methods.
Zhidong Li, Jie Xu 0008, Yang Wang 0002, Glenn Geers, Jun Yang 0033
WACV1
2011 Feature fusion for vehicle detection and tracking with low-angle cameras
abstract
In this paper, we address the problem of vehicle detection and tracking with low-angle cameras by combining windshield detection and feature points clustering, effectively fusing several primitive image features such as color, edge and interest point. By exploring various heterogenous features and multiple vehicle models, we achieve at least two improvements over the existing methods: higher detection accuracy and the ability to distinguish different vehicle types. Our experiments on real-world traffic video sequences demonstrate the benefits of feature fusion and the improved performance.
Jun Yang 0033, Yang Wang 0002, Arcot Sowmya, Zhidong Li, Bang Zhang, Jie Xu 0008
WACV4
2010 Unsupervised Moving Object Detection with On-line Generalized Hough Transform
Jie Xu 0008, Yang Wang 0002, Wei Wang 0011, Jun Yang 0033, Zhidong Li
ACCV (3)5
2010 Spatial-Temporal Affinity Propagation for Feature Clustering with Application to Traffic Video Analysis
Jun Yang 0033, Yang Wang 0002, Arcot Sowmya, Jie Xu 0008, Zhidong Li, Bang Zhang
ACCV (2)5
2010 Affinity Propagation Feature Clustering with Application to Vehicle Detection and Tracking in Road Traffic Surveillance
abstract
In this paper, we investigate the applicability of the newly proposed data clustering method, affinity propagation, in feature points clustering and the task of vehicle detection and tracking in road traffic surveillance. We propose a model-based temporal association scheme and novel preprocessing and postprocessing operations which together with affinity propagation make a quite successful method for the given task. Our experiments demonstrate the effectiveness and efficiency of our method and its superiority over the state-of-the-art algorithm.
Jun Yang 0033, Yang Wang 0002, Arcot Sowmya, Bang Zhang, Jie Xu 0008, Zhidong Li
AVSS6
2010 Image Topic Discovery with Saliency Detection
abstract
This work proposes a biologically inspired approach to integrate latent topic model with saliency detection. Firstly, a saliency detection algorithm is presented to discriminate salient objects from background parts in the image. A hierarchical latent topic model is proposed to discover image topics by combining subtopics of both salient objects and background parts. We test the algorithm on public image datasets for saliency detection and image categorization. The experimental results show that the proposed approach robustly detects salient objects and categorizes image data, and it outperforms state-of-the-art methods for both saliency detection and unsupervised topic modelling.
Zhidong Li, Yang Wang 0002, Jie Xu 0008, John Laird
BMVC1
2010 Saliency based joint topic discovery for object categorization
abstract
We present a novel approach of saliency based image categorization using topic model. In each image, salient foreground objects are discriminated from background scene by saliency detection. Then topic model is used to jointly discover topics of foreground and background. Our approach can categorize images in a completely unsupervised manner and achieve higher performance than previous categorization methods, especially for those images with similar foreground/background.
Zhidong Li, Yang Wang 0002, Glenn Geers, Jun Yang 0033, John Laird
ICIP1
2010 Vehicle detection and tracking with low-angle cameras
abstract
Vision-based vehicle detection is a critical task for traffic monitoring in modern Intelligent Traffic Systems (ITS). Due to the low-angle nature of most traffic surveillance cameras installed in the real world, vehicle detection in such case has to deal with one fundamental challenge — occlusion, which renders most traditional vehicle detection methods ineffective. In this paper, instead of detecting the vehicle as a whole, we propose a vehicle detection algorithm based on windshield model matching. By detecting windshield directly, the algorithm achieves robustness to occlusion. Together with camera calibration and vehicle tracking, the system is able to provide reliable traffic state estimation. Experiments on real traffic videos demonstrate the better performance of our system compared to the state-of-the-art algorithm.
Jun Yang 0033, Yang Wang 0002, Arcot Sowmya, Zhidong Li
ICIP4
2009 Traffic Density Estimation with On-line SVM Classifier
abstract
Information on the vehicular traffic density in an intelligent transport system (ITS) is presently obtained mainly through loop detectors (LD), traffic radars and surveillance cameras. However, the difficulties and cost of installing loop detectors and traffic radars tend to be significant. Currently, a more advanced method of circumventing this is to develop a sort of virtual loop detector (VLD) by using video content understanding technology to simulate behavior of a loop detector and to further estimate the traffic flow from a surveillance camera. Such a virtual loop detector that requires supervised training with human intervention for its setup. Difficulties also arise when attempting to obtain a reliable and real-time VLD under different illumination, weather conditions and static shadows. In this paper, we study the effectiveness of texture features in describing the traffic density, and propose a real-time VLD based on on-line SVM classifier and a background modeling technique (OSVM-BG) to estimate the traffic density information probabilistically and automatically. The system uses feedback from background modeling to train and update its SVM kernel to self-adapt to various lighting environments. Experimental results show that the system outperforms an existing algorithm and achieves an average accuracy of 89.43% under various illumination changes, weather conditions and especially changing static shadows in daytime.
Thanes Wassantachat, Zhidong Li, Yang Wang 0002, Evan Tan
AVSS2
2009 A Machine Learning Framework for Real-Time Traffic Density Detection
abstract
Traffic flow information can be employed in an intelligent transportation system to detect and manage traffic congestion. One of the key elements in determining the traffic flow information is traffic density estimation. The goal of traffic density estimation is to determine the density of vehicles on a given road from loop detectors, traffic radars, or surveillance cameras. However, due to the inflexibility of deploying loop detectors and traffic radars, there is a growing trend of using video-content-understanding technique to determine the traffic flow from a surveillance camera. But difficulties arise when attempting to do this in real-time under changing illumination and weather conditions as well as heavy traffic congestions. In this paper, we attempt to address the problem of real-time traffic density estimation by using a stochastic model called Hidden Markov Models (HMM) to probabilistically determine the traffic density state. Choosing a good set of model parameters for HMMs has a significant impact on the accuracy of traffic density estimation. Thus, we propose a novel feature extraction scheme to represent traffic density, and a novel approach to initialize and construct the HMMs by using an unsupervised clustering technique called AutoClass. We show through extensive experiments that our proposed real-time algorithm achieves an average traffic density estimation accuracy of 96.6% over various different illumination and weather conditions.
Evan Tan, Zhidong Li
Int. J. Pattern Recognit. Artif. Intell.3
2008 An improved mean-shift tracker with kernel prediction and scale optimisation targeting for low-frame-rate video tracking
abstract
The mean-shift (MS) algorithm is widely used in object tracking because of its speed and simplicity. However, it assumes certain overlap of object appearance and smooth change in object scale between consecutive video frames. This assumption is usually violated in a low-frame-rate (LFR) video, which contains fast motion and scale changes. An LFR video is widely adopted in applications such as surveillance systems, where real-time object tracking is highly desirable but the traditional MS algorithm does not perform well. We addressed this problem by proposing a novel and enhanced mean-shift tracker, named SMDShift, that uses kernel prediction and stochastic meta-descent (SMD) optimization method to deal with the kernel position and scale variation when tracking objects in an LFR video. In our experiments, the SMDShift can track fast moving objects with significant scale change in an LFR video sequence on which the traditional mean-shift and Camshift algorithms fail.
Zhidong Li, Nicol N. Schraudolph
ICPR1
2008 Using stochastic gradient-descent scheme in appearance model based face tracking
abstract
Active appearance model (AAM) has been widely used in face tracking and recognition. However, accuracy and efficiency are always two main challenges with the AAM search. The paper therefore proposed a fast appearance-model based 3D face tracking algorithm to track a face appearance with significant translation, rotation, and scaling activities by using stochastic meta-descent (SMD) optimization scheme to accelerate the appearance model search and to improve the tracking efficiency and accuracy. The proposed algorithm constructs an active face appearance model by using several semantic landmark points extracted from each frame and then processes the appearance model search to approximate the model translating, rotating, and scaling by using the SMD filter to minimize the appearance difference between the current model and the new observation. We compared the results with both a conventional AAM and a Camshift filter and found that our algorithm outperforms both two in terms of efficiency and accuracy in tracking a fast moving, rotating, and scaling face object in a video sequence.
Zhidong Li, Adrian Chong, Zhenghua Yu, Nicol N. Schraudolph
MMSP1