Dong Zhang 0002

dblp:68/3245-2 · DBLP profile ↗
← Back
28ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0003-0825-3400ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Security and privacy · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Detecting children with autism spectrum disorder based on script-centric behavior understanding with emotional enhancement
Yueran Pan, Dong Zhang 0002, Hongzhu Deng, Xiaobing Zou, Ming Li 0026
Neurocomputing3
2025 Assessing the Expressive Language Levels of Autistic Children in Home Intervention
abstract
The World Health Organization (WHO) has established the caregiver skill training (CST) program, designed to empower families with children diagnosed with autism spectrum disorder the essential caregiving skills. The joint engagement rating inventory (JERI) protocol evaluates participants’ engagement levels within the CST initiative. Traditionally, rating the expressive language level and use (EXLA) item in JERI relies on retrospective video analysis conducted by qualified professionals, thus incurring substantial labor costs. This study introduces a multimodal behavioral signal-processing framework designed to analyze both child and caregiver behaviors automatically, thereby rating EXLA. Initially, raw audio and video signals are segmented into concise intervals via voice activity detection, speaker diarization and speaker age classification, serving the dual purpose of eliminating nonspeech content and tagging each segment with its respective speaker. Subsequently, we extract an array of audio-visual features, encompassing our proposed interpretable, hand-crafted textual features, end-to-end audio embeddings and end-to-end video embeddings. Finally, these features are fused at the feature level to train a linear regression model aimed at predicting the EXLA scores. Our framework has been evaluated on the largest in-the-wild database currently available under the CST program. Experimental results indicate that the proposed system achieves a Pearson correlation coefficient of 0.768 against the expert ratings, evidencing promising performance comparable to that of human experts.
Yueran Pan, Biyuan Chen, Ming Cheng 0005, Dong Zhang 0002, Hongzhu Deng, Xiaobing Zou, Ming Li 0026
IEEE Trans. Comput. Soc. Syst.5
2024 A dual-branch network based on optical flow learning and semantic consistency for macro-expression spotting
Yun Xian, Dong Zhang 0002, Xingzhi Wang, Dah-Jye Lee
Appl. Intell.2
2024 Implementing the Affective Mechanism for Group Emotion Recognition With a New Graph Convolutional Network Architecture
abstract
Research on social psychology has revealed the existence of an affective mechanism in a human group, which is the group members spread their emotions to one another, the emotions of the group members form the group emotion, and the group emotion as a powerful force shapes the group members' emotions. Current group emotion recognition methods focus on how the emotions of the group members form the group-level emotion but rarely take into account how the group emotion feeds back to the group members instantaneously. This paper proposes a new graph convolutional network architecture to characterize this unique affective mechanism for group emotion recognition. We regard the group members as the nodes of the graph and introduce a pseudo node into the graph to represent the role of the group. This paper uses graph convolutional networks to model the emotional interactions within the group from a static image and constructs an effective emotional representation at the group level for recognition. Experiment results on three widely used datasets for group emotion recognition show that our proposed method achieved superior performance in terms of recognition accuracy compared to the state-of-the-art methods.
Xingzhi Wang, Dong Zhang 0002, Dah-Jye Lee
IEEE Trans. Affect. Comput.2
2024 Joint Training on Multiple Datasets With Inconsistent Labeling Criteria for Facial Expression Recognition
abstract
One potential way to enhance the performance of facial expression recognition (FER) is to augment the training set by increasing the number of samples. By incorporating multiple FER datasets, deep learning models can extract more discriminative features. However, the inconsistent labeling criteria and subjective biases found in annotated FER datasets can significantly hinder the recognition accuracy of deep learning models when handling mixed datasets. Effectively perform joint training on multiple datasets remains a challenging task. In this study, we propose a joint training method for training an FER model using multiple FER datasets. Our method consists of four steps: (1) selecting a subset from the additional dataset, (2) generating pseudo-continuous labels for the target dataset, (3) refining the labels of different datasets using continuous label mapping and discrete label relabeling according to the labeling criteria of the target dataset, and (4) jointly training the model using multi-task learning. We conduct joint training experiments on two popular in-the-wild FER benchmark databases, RAF-DB and CAER-S, while utilizing the AffectNet dataset as an additional dataset. The experimental results demonstrate that our proposed method outperforms the direct merging of different FER datasets into a single training set and achieves state-of-the-art performance on RAF-DB and CAER-S with accuracies of 92.24% and 94.57%, respectively.
Chengyan Yu, Dong Zhang 0002, Ming Li 0026
IEEE Trans. Affect. Comput.2
2023 Human pose estimation for low-resolution image using 1-D heatmaps and offset regression
Cailong Chi, Dong Zhang 0002, Zhesi Zhu, Xingzhi Wang, Dah-Jye Lee
Multim. Tools Appl.2
2023 STCAM: Spatial-Temporal and Channel Attention Module for Dynamic Facial Expression Recognition
abstract
Capturing the dynamics of facial expression progression in video is an essential and challenging task for facial expression recognition (FER). In this article, we propose an effective framework to address this challenge. We develop a C3D-based network architecture, 3D-Inception-ResNet, to extract spatial-temporal features from the dynamic facial expression image sequence. A Spatial-Temporal and Channel Attention Module (STCAM) is proposed to explicitly exploit the holistic spatial-temporal and channel-wise correlations among the extracted features. Specifically, the proposed STCAM calculates a channel-wise and a spatial-temporal-wise attention map to enhance the features along the corresponding feature dimensions for more representative features. We evaluate our method on three popular dynamic facial expression recognition datasets, CK+, Oulu-CASIA, and MMI. Experimental results show that our method achieves better or comparable performance compared to the state-of-the-art approaches.
Weicong Chen 0003, Dong Zhang 0002, Ming Li 0026, Dah-Jye Lee
IEEE Trans. Affect. Comput.2
2023 Computer-Aided Autism Spectrum Disorder Diagnosis With Behavior Signal Processing
abstract
Behavioral observation plays an essential role in the diagnosis of Autism Spectrum Disorder (ASD) by analyzing children's atypical patterns in social activities (e.g., impaired social interaction, restricted interests, and repetitive behavior). To date, this process still heavily relies on the questionnaire survey, clinical observation, or retrospective video analysis, leading to high demand for professionals with massive labor costs. This article proposes a standardized platform for stimulating, gathering, analyzing, modeling, and interpreting human behavioral data in the application of computer-aided ASD diagnosis. By a structured assessment process, the proposed system can automatically evaluate children's multiple social interaction skills using the captured audio-visual data and provide the final diagnostic suggestions. We collect a multimodal behavioral database of 95 participants (71 children with ASD and 24 age-matched typical controls) in a real clinic environment, the Third Affiliated Hospital of Sun Yat-sen University, China. On the clinical database, our proposed computer-aided ASD diagnosis system obtains an accuracy of 88.42% for identifying ASD children with an average age of 24 months, representing a performance comparable to top-level human experts. As a unified and replicable solution, it has good potential to be promoted to less developed areas with limited high-quality medical resources.
Ming Cheng 0005, Yixiang Xie, Yueran Pan, Xiao Li 0048, Chengyan Yu, Dong Zhang 0002, Xiaoqian Huang, Cong You, Yuanyuan Zou 0003, Yuchong Liu, Fengjing Liang, Huilin Zhu, Chun Tang, Hongzhu Deng, Xiaobing Zou, Ming Li 0026
IEEE Trans. Affect. Comput.8
2023 Typical Facial Expression Network Using a Facial Feature Decoupler and Spatial-Temporal Learning
abstract
Facial expression recognition (FER) accuracy is often affected by an individual’s unique facial characteristics. Recognition performance can be improved if the influence from these physical characteristics is minimized. Using video instead of single image for FER provides better results but requires extracting temporal features and the spatial structure of facial expressions in an integrated manner. We propose a new network called Typical Facial Expression Network (TFEN) to address both challenges. TFEN uses two deep two-dimensional (2D) convolutional neural networks (CNNs) to extract facial and expression features from input video. A facial feature decoupler decouples facial features from expression features to minimize the influence from inter-subject face variations. These networks combine with a 3D CNN and form a spatial-temporal learning network to jointly explore the spatial-temporal features in a video. A facial recognition network works as an adversarial network to refine the facial feature decoupler and the network performance by minimizing the residual influence of facial features after decoupling. The whole network is trained with an adversarial algorithm to improve FER performance. TFEN was evaluated on four popular dynamic FER datasets. Experimental results show TFEN achieves or outperforms the recognition accuracy of state-of-the-art approaches.
Jianing Teng, Dong Zhang 0002, Ming Li 0026, Dah-Jye Lee
IEEE Trans. Affect. Comput.2
2023 A Self-Fusion Network Based on Contrastive Learning for Group Emotion Recognition
abstract
Group emotion recognition (GER) from image has attracted much attention in recent years. Networks using attention mechanism for GER have shown great potential. However, the performance of the current attention-based GER networks suffers from the indistinctive features of individuals in the group, poor feature fusion weights, and the lack of semantic information of the objects in the image. We present a new framework that is composed of three networks, FacesNet, SceneNet, and ObjectsNet, to address these shortcomings. This new framework is designed to recognize group emotion by exploiting the information from the faces, scene, and objects in image. In FacesNet, we use contrastive learning to help the network extract distinctive emotion features and a new attention mechanism named self-fusion module to generate precise fusion weights for aggregation of individual facial features. We design SceneNet to capture the multiscale scene features to exploit the emotion cues from the scene. We construct a fully connected network named ObjectsNet to classify the semantic features of the objects. Finally, we linearly integrate the outputs of these three networks as the final output of this unique framework for GER. Experiment results on three datasets for GER show that our proposed framework achieved better performance in terms of recognition accuracy compared with the state-of-the-art methods.
Xingzhi Wang, Dong Zhang 0002, Hongzhou Tan, Dah-Jye Lee
IEEE Trans. Comput. Soc. Syst.2
2023 Accurate Head Pose Estimation Using Image Rectification and a Lightweight Convolutional Neural Network
abstract
Head pose estimation is an important step for many human-computer interaction applications such as face detection, facial recognition, and facial expression classification. Accurate head pose estimation benefits these applications that require face images as the input. Most head pose estimation methods suffer from perspective distortion because the users do not always align their face perfectly with the camera. This paper presents a new approach that uses image rectification to reduce the negative effect of perspective distortion and a lightweight convolutional neural network to obtain highly accurate head pose estimations. The proposed method calculates the angle between the optical axis of the camera and the projection vector of the center of the face. The face image is rectified using this estimated angle through perspective transformation. A lightweight network that is only 0.88 MB in size is designed to take the rectified face image as the input to perform head pose estimation. The output of the network, the head pose estimation of the rectified face image, is transformed back to the camera coordinate system as the final head pose estimation. Experiments on public benchmark datasets show that the proposed image rectification method and the newly designed lightweight network improve the accuracy of head pose estimation remarkably. Compared with state-of-the-art methods, our approach achieves both higher accuracy and faster processing speed.
Xiao Li 0048, Dong Zhang 0002, Ming Li 0026, Dah-Jye Lee
IEEE Trans. Multim.2
2022 A new multi-feature fusion based convolutional neural network for facial expression recognition
Dong Zhang 0002, Dah-Jye Lee
Appl. Intell.2
2021 The 2020 Personalized Voice Trigger Challenge: Open Datasets, Evaluation Metrics, Baseline System and Results
Xingming Wang, Xiaoyi Qin, Yinping Zhang, Junjie Wang 0010, Dong Zhang 0002, Ming Li 0026
Interspeech7
2019 Automatic fabric defect detection with a wide-and-compact network
Dong Zhang 0002, Dah-Jye Lee
Neurocomputing2
2019 IIRNet: A lightweight deep neural network using intensely inverted residuals for image recognition
Dong Zhang 0002, Dah-Jye Lee
Image Vis. Comput.2
2019 Recognition of Chinese food using convolutional neural network
Jianing Teng, Dong Zhang 0002, Dah-Jye Lee, Yao Chou
Multim. Tools Appl.2
2017 A parallel convolutional neural network architecture for stereo vision estimation
abstract
Extracting depth information from the stereo image pair is a commonly used method in 3-D computer vision. For robotics and unmanned vehicle applications that require real-time performance, speed is often more important than accuracy. In recent years, Convolutional Neural Networks (CNNs) have shown great success in many computer vision applications including classification, segmentation, object detection, edge detection, and stereo vision estimation. Existing network architectures for stereo vision estimation predict very little information during the forward pass and are only able to calculate the disparity for one pixel at a time. In this paper, we propose a parallel architecture to speed up disparity map computation by simultaneously processing all pixels on one horizontal line. We train and test our network on five Middlebury datasets. Our parallel architecture achieves at least a 10x speedup compared to existing networks. Its accuracy is also very competitive.
Yao Chou, Dah-Jye Lee, Dong Zhang 0002, Karina Hill
ICIP3
2017 Edge Detection Using Convolutional Neural Networks for Nematode Development and Adaptation Analysis
Yao Chou, Dah-Jye Lee, Dong Zhang 0002
ICVS3
2017 A new computer vision based multi-indentation inspection system for ceramics
YingIyad Jafar Liu, Dong Zhang 0002, Huacang Peng, Yonghong Zhu
Multim. Tools Appl.3
2015 Automatic fish taxonomy using evolution-constructed features for invasive species removal
Dong Zhang 0002, Kirt D. Lillywhite, Dah-Jye Lee, Beau J. Tippetts
Pattern Anal. Appl.1
2014 Steganalysis based on distribution characters of stego-images in reduced dimension space
Guoming Chen 0001, Dong Zhang 0002, Duanning Zhou
Multim. Tools Appl.3
2014 Seeing Eye Phone: a smart phone-based indoor localization and guidance system for the visually impaired
Dong Zhang 0002, Dah-Jye Lee, Brandon Taylor
Mach. Vis. Appl.1
2012 An efficient shape analysis method for shrimp quality evaluation
abstract
Two grading criteria used in determining shrimp product quality and value by the shrimp industry are: 1. Presence or percentage of black spot, measured as a percentage of the total body surface. 2. Shape quality referring to whole shrimp and broken pieces. Black spots (melanoma) on the shrimp surface are evidence of aging shrimp and are considered defects that must be removed from the main production line. Shape quality is measured as the size and the completeness of the body. Broken shrimp pieces are considered a product defect and also must be removed from the main production line. Black spot detection is a simple task for a well-designed machine vision system, which provides consistent and controlled lighting. Shape analysis, on the other hand, is a challenging task because it involves contour extraction and shape analysis. In this paper, a simple, fast, and accurate shape analysis method using Turn Angle Cross-correlation is developed for shrimp quality evaluation. Our analysis results validate that the performance of the proposed shape analysis method is suitable for real-time inspection for commercial applications.
Dah-Jye Lee, Guangming Xiong, Robert M. Lane, Dong Zhang 0002
ICARCV4
2012 Smart phone-based Indoor guidance system for the visually impaired
abstract
A smart phone camera based indoor guidance system to aid the visually impaired is presented. Most proposed systems for aiding the visually impaired with indoor navigation are not feasible for widespread use due to cost, usability, or portability. A smart phone vision- based indoor guidance system that is simple, accessible, inexpensive, and discrete is developed to aid the visually impaired to navigate unfamiliar environments such as public buildings. The system consists of a smart phone and a server. The smart phone captures and transmits pictures of the user's surroundings to the server. The server processes the images and matches them to a database of stored images of the building. After matching features, the location and orientation of the person is calculated using 3D location correspondence data stored for features of each image. Positional information is then transmitted back to the smart phone and communicated to the user via text-to-speech. This paper focuses on developing the vision technology for this unique application rather than building the complete system. Experimental results demonstrate the ability of the system to quickly and accurately determine the pose of the user in a university building.
Brandon Taylor, Dah-Jye Lee, Dong Zhang 0002, Guangming Xiong
ICARCV3
2012 Security of cass data hiding scheme under the scenarios of KMA and WOA
abstract
This paper presents a theoretical analysis on the security of CASS (Correlation-and-bit-Aware Spread-Spectrum) data hiding scheme for the first time. By evaluated with the residual entropy of the secret key and the mutual information between the observations and the secret key, the security of CASS is investigated under the scenarios of Known Message Attack and Watermarked Only Attack. In addition, this paper compares data hiding security between CASS and the conventional Additive Spread-Spectrum (Add-SS) data hiding scheme. Theoretical analysis and simulation results show CASS scheme outperforms Add-SS scheme in terms of data hiding security when Document to Watermark Ratio (DWR) is low and performs comparably when DWR is high.
Dong Zhang 0002, Dah-Jye Lee
ICASSP1
2009 Lessons learned in developing a low-cost high performance medical imaging cluster
abstract
This paper explores the usefulness of the Sony PlayStation 3reg(PS3) for medical image processing. Medical image processing often entails dealing with a large number of high resolution images, requiring a large amount of computational power to process. The PS3 is powered by the cell broadband engine, a microprocessor created by IBM, capable of rapid numeric computation with low power requirements that has helped the unit to become a popular gaming unit. The unit can be repurposed as a low-cost high performance computing platform. In order to demonstrate the computational abilities of the PS3, several basic image processing tasks are implemented and compared with desktop PCs, equipped with general-purpose microprocessors. This article describes lessons learned in the process of building some fundamental image processing tasks. The article also describes the architecture of the cell broadband engine and provides information about developing applications given the architecture. The article also provides an introduction to setting up a high-performance image processing environment with a cluster of such relatively inexpensive PS3 units.
Kirt D. Lillywhite, Dah-Jye Lee, Sameer K. Antani, Dong Zhang 0002, L. Rodney Long
CBMS4
2009 Real-time human detection using histograms of oriented gradients on a GPU
abstract
Human detection has always been an important part of computer vision but many implementations lack the real-time performance that real world applications require. This paper presents a real-time implementation of human detection in video using the state-of-the-art histograms of oriented gradients method. Each image in the video sequence is tested at multiple scales using a sliding window. Histograms of oriented gradients are created for each window and passed to a support vector machine to classify it as human or not. The histograms of oriented gradients method is implemented on a GPU using the NVIDIA CUDA architecture. The implementation significantly speeds up computation, achieving approximately 38 frames a second on VGA video while testing 11,160 windows per frame. Accuracy remains comparable to the CPU implementation. The flexibility and computational power the GPU affords users is discussed. These discussions should benefit those researchers who are interested in using a GPU for high-performance computing tasks.
Kirt D. Lillywhite, Dah-Jye Lee, Dong Zhang 0002
WACV3
2008 GSM Based Security Analysis for Add-SS Watermarking
Dong Zhang 0002, Jiangqun Ni, Dah-Jye Lee, Jiwu Huang
IWDW1