EDBT 2026 Demo / reviewers in the wild / expert
Yuzhu Ji
dblp:169/2012
· DBLP profile ↗
41ranked-venue papers
9as first author
24since 2021 · last 2026
0000-0003-3589-3884ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Systems, architecture and hardware · 6Databases, data management, data science and information retrieval · 6 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | One-Shot Federated Clustering of Non-Independent Completely Distributed DataabstractFederated Learning (FL) that extracts data knowledge while protecting the privacy of multiple clients has achieved remarkable results in distributed privacy-preserving IoT systems, including smart traffic flow monitoring, smart grid load balancing, and so on. Since most data collected from edge devices are unlabeled, unsupervised Federated Clustering (FC) is becoming increasingly popular for exploring pattern knowledge from complex distributed data. However, due to the lack of label guidance, the common Non-Independent and Identically Distributed (Non-IID) issue of clients have greatly challenged FC by posing the following problems: How to fuse pattern knowledge (i.e., cluster distribution) from Non-IID clients; How are the cluster distributions among clients related; and How does this relationship connect with the global knowledge fusion? In this paper, a more tricky but overlooked phenomenon in Non-IID is revealed, which bottlenecks the clustering performance of the existing FC approaches. That is, different clients could fragment a cluster, and accordingly, a more generalized Non-IID concept, i.e., Non-ICD (Non-Independent Completely Distributed), is derived. To tackle the above FC challenges, a new framework named GOLD (Global Oriented Local Distribution Learning) is proposed. GOLD first finely explores the potential incomplete local cluster distributions of clients, then uploads the distribution summarization to the server for global fusion, and finally performs local cluster enhancement under the guidance of the global distribution. Extensive experiments, including significance tests, ablation studies, scalability evaluations, qualitative results, etc., have been conducted to show the superiority of GOLD. Yiqun Zhang 0006, Shenghong Cai, Zihua Yang, Sen Feng, Yuzhu Ji, Haijun Zhang 0002 |
IEEE Internet Things J. | 5 |
| 2026 | An End-to-end learning for classification and segmentation of breast cancer
Jinling Chen, Zhanbo Tan, Zhuowei Tang, Jihong Wei, Qi Ke, Yuzhu Ji, Ziqing Gao |
Multim. Tools Appl. | 7 |
| 2026 | Online Heterogeneous Feature SelectionabstractMany real-world datasets contain high-dimensional heterogeneous features, exhibiting complex and evolving distributions. The coexistence of high dimensionality and heterogeneity poses challenges for reliable feature selection and real-time analysis, while most existing feature selection solutions either assume that the features are of the same type or struggle to handle extremely high-dimensional features. Moreover, these methods are usually designed for static datasets, neglecting the dynamic capture of heterogeneous interfeature relationships in real-time environments. To address these challenges, we propose a new feature selection method called graph-unified adaptive decision boundary enhancement (GRADE) for online heterogeneous feature selection (OHFS). To provide a reliable foundation for evaluating feature subsets under dynamic and heterogeneous data streams, an incremental graph-unified metric (IGUM) is introduced. It mitigates information loss between heterogeneous features by leveraging graph structures to unify feature-value-level and interfeature-level relationships. With such a consistent relation measure, an adaptive density-guided neighborhood relation (ADNR) is proposed to assess the capability of selected feature subsets to classify samples. Since it dynamically captures prominent neighborhood regions, local decision boundaries can thus be precisely delineated. It turns out that GRADE can obtain a more concise feature subset while achieving competitive classification accuracy. Besides, GRADE is parameter-free and very efficient compared with state-of-the-art methods. Comprehensive experimental evaluations, including significance tests, ablation studies, efficiency evaluation, and case studies, have been conducted to verify the efficacy of GRADE. Yiqun Zhang 0006, Xinxi Chen, Lang Zhao, Yuzhu Ji, Peng Liu 0045, Yiu-Ming Cheung |
IEEE Trans. Cybern. | 4 |
| 2025 | Robust Qualitative Data Clustering via Learnable Multi-Metric Space FusionabstractUnderstanding categorical data with vague qualitative values by forming clusters is crucial in many data-driven AI fields. Compared with numerical data with its quantitative values embedded in well-defined Euclidean distance space, distances of the qualitative values are naturally unknown and are specially defined for certain data types or tasks. This paper, therefore, proposes a distance metric space fusion framework, which learns to fuse multiple distance metrics to form a statistical information-complete and prior knowledge-comprehensive metric for robust and accurate cluster analysis of qualitative data. To better serve various clustering tasks, the metric fusion objective is incorporated into the clustering objective through iterative learning. It turns out that the proposed method stably demonstrates superiority on various challenging real benchmark datasets. Extensive experiments including significance tests, ablation studies, etc. validate its efficacy. Source code of the proposed method is available at https://github.com/Sen-Feng/ICASSP-MSF/tree/main/CODE. Sen Feng, Mingjie Zhao 0003, Zhanpei Huang, Yuzhu Ji, Yiqun Zhang 0006, Yiu-Ming Cheung |
ICASSP | 4 |
| 2025 | SmartNet: One-shot Talking Head Synthesis via Subtle Motion and Appearance CompensationabstractOne-shot talking head synthesis aims to animate a source person’s portrait with driving video sequences. Recent facial keypoint-based methods have achieved remarkable animation performance and produced high-quality results. However, it remains challenging to perform cross-identity face reenactment by transferring subtle facial motions with correct geometry and appearance. To break the above limitations, in this paper, we propose a subtle motion compensation network to recover correct facial expressions by leveraging the decoupled 3D Morphable Model (3DMM) coefficient. In addition, to generate faithful animation results, a facial appearance feature memory bank is designed to learn accurate facial features and better recover the appearance. Experimental results have demonstrated that our proposed model can outperform state-of-the-art methods by generating faithful videos with correct subtle motion transfer and consistent identity preserving. Yuzhu Ji, An Zeng, Dan Pan 0001, Yiqun Zhang 0006, Haijun Zhang 0002 |
ICASSP | 2 |
| 2025 | Coarse-To-Fine Graph Reasoning for 3D Hand Mesh Reconstructionabstract3D hand mesh reconstruction from 2D images is crucial for various computer visual tasks such as virtual reality and human-computer interaction, while it remains a challenging problem due to changed hand poses and diverse self-/cross-hand occlusions. In this paper, we propose a graph-based reasoning network for 3D hand mesh reconstruction from a single 2D RGB image by recovering the fine-grained 3D hand mesh keypoints in a coarse-to-fine manner. First, we extract the hand-joint-specific features to capture the overall hand structure via a CNN backbone with a linear projection sampling operation. Based on the hand-joint locations, we further progressively learn more hand mesh keypoints to construct the coarse hand shape via hierarchical attention-embedded graph learning layers. Finally, we leverage the 2D shallow semantic features to refine the coarse hand mesh keypoints into fine-grained 3D hand mesh vertices coordinates via cascaded graph learning layer with linear mapping. Experimental results on the widely used hand databases show that our method achieves outstanding performance in both single-hand and two-interactive-hand 3D mesh reconstruction. Dan Fu, Wai Keung Wong, Lunke Fei, Tingting Chai, Yuzhu Ji, Qinghua Zhu 0001 |
ICME | 5 |
| 2025 | SeqPose: An End-to-End Framework to Unify Single-frame and Video-based RGB Category-Level Pose EstimationabstractCategory-level object pose estimation is a longstanding and fundamental task crucial for augmented reality and robotic manipulation applications. Existing RGB-based approaches struggle with multi-stage settings and heavily rely on off-the-shelf techniques, such as object detectors, depth estimators, non-differentiable NOCS shape alignment, etc. Extra dependencies lead to the accumulation of errors and complicate the whole pipeline, limiting the deployment of these approaches in practical applications. This paper streamlined an end-to-end framework unifying the single-frame and video-based category-level pose estimation. Specifically, instead of explicitly introducing extra dependencies, the DINOv2 encoder and depth decoder, as robust semantic and geometric prior extractors, are leveraged to produce intra-frame hierarchical semantic and geometric features. A spatial-temporal sparse query network is developed to model the implicit correspondence and inter-frame correlations between a set of implicit 3D query anchors and intra-frame features. Finally, a pose prediction head is employed using the bipartite matching algorithm. Experimental results demonstrate that our model achieves state-of-the-art performance compared with RGB-based categorical pose estimation methods on the REAL275 and CAMERA25 datasets. Our code is available at https://andrewchiyz.github.io/vision.3dv.seqpose/. Yuzhu Ji, Mingshan Sun, Jianyang Shi, Xiaoke Jiang, Yiqun Zhang 0006, Haijun Zhang 0002 |
IJCAI | 1 |
| 2025 | Towards Clustering of Incomplete Mixed-Attribute DataabstractABSTRACT Clustering analysis is one of the most important data mining and knowledge discovery tools in real applications. Since the widespread presence of missing values hampers clustering performance, missing values imputation becomes necessary for data pre‐processing. However, for the common datasets composed of both numerical and categorical attributes (also known as mixed‐attribute datasets), most existing imputation methods suffer from the following three limitations: (1) Only feasible for a certain type of attribute; (2) Encounter difficulties in considering the interdependence between different types of attributes; (3) Short in exploiting the information provided by the incomplete mix‐valued objects. As a result, the original data distribution can be ill‐restored, misleading the downstream clustering tasks. This paper therefore proposes a clustering‐imputation co‐learning method for incomplete mixed‐attribute datasets to address these issues. This method integrates imputation and clustering into one learning process, emphasising the interrelationships between mixed attributes during the imputation process and exploiting the information of incomplete objectsduring clustering. It turns out that appropriate recovery of the dataset and accurate clustering can be better achieved through a cross‐coupling manner. Experiments on various datasets validate the promising efficacy of the proposed method. Chuyao Zhang, Xinxi Chen, Zexi Tan, Fangqing Gu, Yuzhu Ji, Yiqun Zhang 0006 |
Expert Syst. J. Knowl. Eng. | 5 |
| 2025 | One-Shot Human Motion Transfer via Occlusion-Robust Flow Prediction and Neural TexturingabstractHuman motion transfer aims at animating a static source image with a driving video. While recent advances in one-shot human motion transfer have led to significant improvement in results, it remains challenging for methods with 2D body landmarks, skeleton and semantic mask to accurately capture correspondences between source and driving poses due to the large variation in motion and articulation complexity. In addition, the accuracy and precision of DensePose degrade the image quality for neural-rendering-based methods. To address the limitations and by both considering the importance of appearance and geometry for motion transfer, in this work, we proposed a unified framework that combines multi-scale feature warping and neural texture mapping to recover better 2D appearance and 2.5D geometry, partly by exploiting the information from DensePose, yet adapting to its inherent limited accuracy. Our model takes advantage of multiple modalities by jointly training and fusing them, which allows it to robust neural texture features that cope with geometric errors as well as multi-scale dense motion flow that better preserves appearance. Experimental results with full and half-view body video datasets demonstrate that our model can generalize well and achieve competitive results, and that it is particularly effective in handling challenging cases such as those with substantial self-occlusions. Yuzhu Ji, Chuanxia Zheng, Tat-Jen Cham |
IEEE Trans. Multim. | 1 |
| 2025 | Learning Self-Growth Maps for Fast and Accurate Imbalanced Streaming Data ClusteringabstractStreaming data clustering is a popular research topic in data mining and machine learning. Since streaming data is usually analyzed in data chunks, it is more susceptible to encountering the dynamic cluster imbalance issue. That is, the imbalance ratio (IR) of clusters changes over time, which can easily lead to fluctuations in either the accuracy or the efficiency of streaming data clustering. Therefore, an accurate and efficient streaming data clustering approach is proposed to adapt to the drifting and imbalanced cluster distributions. We first design a self-growth map (SGM) that can automatically arrange neurons on demand according to local distribution, and thus achieve fast and incremental adaptation to the streaming distributions. Since SGM allocates an excess number of density-sensitive neurons to describe the global distribution, it can avoid missing small clusters among imbalanced distributions. We also propose a fast hierarchical merging (HM) strategy to combine the neurons that break up the relatively large clusters. It exploits the maintained SGM to quickly retrieve the intracluster distribution pairs for merging, which circumvents the most laborious global searching. It turns out that the proposed SGM can incrementally adapt to the distributions of new chunks, and the self-growth map-guided hierarchical merging for the imbalanced data clustering (SOHI) approach can quickly explore a true number of imbalanced clusters. Extensive experiments demonstrate that SOHI can efficiently and accurately explore cluster distributions for streaming data. Yiqun Zhang 0006, Sen Feng, Zexi Tan, Xiaopeng Luo, Yuzhu Ji, Yiu-Ming Cheung |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | MGOD: Multi-Granular Outlier Detection with Clustlier AnalysisabstractUnsupervised Outlier Detection (UOD) is crucial for the analysis of biomedical and health data with undesirable outliers. However, the complex distribution of real data often brings difficulties to UOD where the "masking effect", i.e., only a small number of densely distributed outliers (also called clustliers) can collectively mask themselves from being detected, is particularly challenging. Another difficulty derived from this is how to distinguish clustliers from small clusters. Therefore, we propose a novel Multi-Granular Outlier Detector (MGOD). It first partitions the dataset into subsets with natural neighbor topological relationships to circumvent the non-trivial neighbor range setting. Then it effectively detects both clustliers and isolated samples (also called scatliers) based on a newly designed anomaly score. The score comprehensively takes into account the density and connectivity of samples to reflect different extents and types of abnormality. It turns out that MGOD is accurate and highly interpretable. The performance of MGOD is also robust to the involved hyper-parameters, which are easy to set. Comprehensive evaluations have been conducted to compare seven counterparts on 15 datasets, most of which are biomedical datasets. The results of significance tests confirm the effectiveness and superiority of MGOD. The source code is opened at https://anonymous.4open.science/r/MGOD-C531. Qingsheng Chen, Mingjie Zhao 0003, Yuzhu Ji, Xiaopeng Luo, Yiqun Zhang 0006, Yue Zhang 0045 |
BIBM | 3 |
| 2024 | Efficient Topology-Driven Clustering for Imbalanced Streaming Biomedical Data AnalysisabstractClustering drifting data is common in the field of biomedical data analysis. Data chunks collected at different periods often exhibit clusters with significantly different sizes, and drifting distributions of clusters also appear frequently. We call such composite phenomenon imbalance-drifting, which can severely impact the accuracy and efficiency of cluster analysis. Therefore, we propose a topology-representation-based clustering paradigm, which first learns an informative global data representation in a self-organizing manner to obtain a map with nested representative data points. Then fast and accurate clustering is facilitated by quickly retrieving similar data points according to the topology. As the constructed Self-Organizing Map (SOM) is exploited for informative representation, micro partition, and quick merging, to achieve advanced clustering under imbalance-drifting, the proposed approach is thus called Tri-Squeezing SOM for Clustering (TSSC). It turns out that TSSC significantly reduces the time complexity for clustering an n-scale imbalance-streaming data without sacrificing accuracy. Moreover, TSSC can automatically determine the number of clusters k, and features interpretability and hyper-parameter robustness. Extensive results on both biomedical datasets and synthetic datasets verify the superiority of TSSC. Xiaopeng Luo, Yiqun Zhang 0006, Yuzhu Ji, Peng Liu 0045, Taoting Xiao |
BIBM | 3 |
| 2024 | CPNet: 3D Semantic Relation and Geometry Context Prior Network for Multi-Organ SegmentationabstractAutomatic multi-organ segmentation of the abdominal region is a critical yet challenging task in computer-aided medical image analysis. Recent advances in CNN- and Transformer-based encoder-decoder models tend to implicitly learn context features by using enhanced effective receptive fields to capture local and global range dependencies. However, due to the complex anatomical structure, those models cannot recover the anatomical topology properly and result in broken organs with inaccurate semantic labels. Therefore, in this paper, by considering the anatomy priors of multi-organs, we propose a Context Prior Network, namely CPNet, which integrates the 3D context semantic relations and geometry priors as explicit anatomical constraints. Specifically, a Semantic Relation Prior Propagation (SRPP) module is designed to propagate the semantic relations between voxels progressively. Moreover, a Multiple Context Prior Prediction (MCPP) module is adopted to preserve the accurate shape and topology by recovering 3D contours and surface normal. Experimental results demonstrate our proposed model outperforms state-of-the-art models for multi-organ segmentation on Abdomen CT and MRI datasets, especially for recovering organs with correct semantic labels and anatomical structures. Yuzhu Ji, Mingshan Sun, Yiqun Zhang 0006, Haijun Zhang 0002 |
ECAI | 1 |
| 2024 | QGRL: Quaternion Graph Representation Learning for Heterogeneous Feature Data ClusteringabstractClustering is one of the most commonly used techniques for unsupervised data analysis.As real data sets are usually composed of numerical and categorical features that are heterogeneous in nature, the heterogeneity in the distance metric and feature coupling prevents deep representation learning from achieving satisfactory clustering accuracy.Currently, supervised Quaternion Representation Learning (QRL) has achieved remarkable success in efficiently learning informative representations of coupled features from multiple views derived endogenously from the original data.To inherit the advantages of QRL for unsupervised heterogeneous feature representation learning, we propose a deep QRL model that works in an encoder-decoder manner.To ensure that the implicit couplings of heterogeneous feature data can be well characterized by representation learning, a hierarchical coupling encoding strategy is designed to convert the data set into an attributed graph to be the input of QRL.We also integrate the clustering objective into the model training to facilitate a joint optimization of the representation and clustering.Extensive experimental evaluations illustrate the superiority of the proposed Quaternion Graph Representation Learning (QGRL) method in terms of clustering accuracy and robustness to various data sets composed of arbitrary combinations of numerical and categorical features.The source code is opened at https://github.com/Juny-Chen/QGRL.git. Yuzhu Ji, Yiqun Zhang 0006, Yiu-Ming Cheung |
KDD | 2 |
| 2023 | Time-Series Data Imputation via Realistic Masking-Guided Tri-Attention Bi-GRUabstractTime series data with missing values are ubiquitous in real applications due to various unforeseen faults during data generation, storage, and transmission. Time-Series Data Imputation (TSDI) is thus crucial to many temporal data analysis tasks. However, existing works usually consider only one of the following two issues: (1) intra-feature temporal dependency, and (2) inter-feature correlation, leading to the overlook of complex coupling information in imputation. To achieve more accurate TDSI, we design a novel imputation model called TABiG, which delicately preserves the short-term, long-term, and inter-feature dependencies by attention mechanisms in a delay error-reduced bi-directional architecture. That is, it leverages GRU to model short-term temporal dependencies and adopts self-attention mechanisms hierarchically to capture long-term temporal dependencies and inter-feature correlations. The multiple self-attention mechanisms are nested in a bi-directional structure to alleviate the problem of delay errors in RNN-like structures. To facilitate model training with higher generalization, a masking strategy that mimics various extreme real missing situations beyond the simple random ones has been adopted for generating self-supervised learning tasks. Comprehensive experiments demonstrate that TABiG significantly outperforms most state-of-the-art imputation counterparts. Complementary results and source code can be accessed at https://github.com/Zhang2112105189/TABiG Yiqun Zhang 0006, An Zeng, Dan Pan 0001, Yuzhu Ji |
ECAI | 5 |
| 2023 | CFNet: A Coarse-to-Fine Framework for Coronary Artery Segmentation
Shiting He, Yuzhu Ji, Yiqun Zhang 0006, An Zeng, Dan Pan 0001 |
PRCV (5) | 2 |
| 2023 | Unsupervised Concept Drift Detection via Imbalanced Cluster Discriminator Learning
Mingjie Zhao 0003, Yiqun Zhang 0006, Yuzhu Ji, Yang Lu 0009 |
PRCV (3) | 3 |
| 2022 | Heterogeneous Drift Learning: Classification of Mix-Attribute Data with Concept DriftsabstractAs many real data sets (e.g., social, financial, and medical data sets) are successively generated in evolution with the ever-changing environment, classification for data stream with concept drift attracts increasing attention in the fields of machine learning and data mining. However, to the best of our knowledge, existing works mainly consider the concept drift issue while ignoring another common characteristic of real data, i.e., existence of awkward heterogeneity caused by mixture of numerical and categorical attributes. It is worth noting that tackling both the concept drift and heterogeneity problems together is exponentially more challenging than dealing with only one of them. This paper, therefore, proposes an ensemble learning approach for the classification of numerical-and-categorical-attribute data (also called mixed data hereinafter) under concept drift. We first design a unified metric to appropriately address the heterogeneity of numerical and categorical attributes. Then a base classifier that can appropriately fuse the information provided by the heterogeneous attributes is formed accordingly. Furthermore, to make the classification adapt to the complex concept drifts demonstrated on the heterogeneous attributes, two types of base classifier ensembles are dynamically learned on the fly. Experimental results on various real mixed data sets with concept drifts demonstrate the efficacy of the proposed method. Lang Zhao, Yiqun Zhang 0006, Yuzhu Ji, An Zeng, Fangqing Gu, Xiaopeng Luo |
DSAA | 3 |
| 2022 | LGCNet: A local-to-global context-aware feature augmentation network for salient object detection
Yuzhu Ji, Haijun Zhang 0002, Feng Gao 0015, Haofei Sun, Haokun Wei |
Inf. Sci. | 1 |
| 2021 | EWNet: An early warning classification framework for smart grid based on local-to-global perception
Feng Gao 0015, Qun Li 0011, Yuzhu Ji, Shengchang Ji, Haofei Sun, Simeng Feng, Haokun Wei, Haijun Zhang 0002 |
Neurocomputing | 3 |
| 2021 | An empirical study of multi-scale object detection in high resolution UAV images
Haijun Zhang 0002, Mingshan Sun, Qun Li 0011, Ming Liu 0014, Yuzhu Ji |
Neurocomputing | 6 |
| 2021 | CNN-based encoder-decoder networks for salient object detection: A comprehensive review and recent advances
Yuzhu Ji, Haijun Zhang 0002, Zhao Zhang 0001, Ming Liu 0014 |
Inf. Sci. | 1 |
| 2021 | ID-Net: an improved mask R-CNN model for intrusion detection under power grid surveillance
Feng Gao 0015, Shengchang Ji, Qun Li 0011, Yuzhu Ji, Simeng Feng, Haokun Wei |
Neural Comput. Appl. | 5 |
| 2021 | CASNet: A Cross-Attention Siamese Network for Video Salient Object DetectionabstractRecent works on video salient object detection have demonstrated that directly transferring the generalization ability of image-based models to video data without modeling spatial-temporal information remains nontrivial and challenging. Considering both intraframe accuracy and interframe consistency of saliency detection, this article presents a novel cross-attention based encoder-decoder model under the Siamese framework (CASNet) for video salient object detection. A baseline encoder-decoder model trained with Lovász softmax loss function is adopted as a backbone network to guarantee the accuracy of intraframe salient object detection. Self- and cross-attention modules are incorporated into our model in order to preserve the saliency correlation and improve intraframe salient detection consistency. Extensive experimental results obtained by ablation analysis and cross-data set validation demonstrate the effectiveness of our proposed method. Quantitative results indicate that our CASNet model outperforms 19 state-of-the-art image- and video-based methods on six benchmark data sets. Yuzhu Ji, Haijun Zhang 0002, Zequn Jie, Lin Ma 0002, Q. M. Jonathan Wu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Clothescounter: A framework for star-oriented clothes mining from videos
Haijun Zhang 0002, Yuzhu Ji, Q. M. Jonathan Wu |
Neurocomputing | 4 |
| 2020 | Crowd counting by using multi-level density-based spatial information: A Multi-scale CNN framework
Li Dong 0011, Haijun Zhang 0002, Yuzhu Ji |
Inf. Sci. | 3 |
| 2020 | Toward New Retail: A Benchmark Dataset for Smart Unmanned Vending MachinesabstractDeep learning is a popular direction in computer vision and digital image processing. It is widely utilized in many fields, such as robot navigation, intelligent video surveillance, industrial inspection, and aerospace. With the extensive use of deep learning techniques, classification and object detection algorithms have been rapidly developed. In recent years, with the introduction of the concept of “unmanned retail,” object detection, and image classification play a central role in unmanned retail applications. However, open-source datasets of traditional classification and object detection have not yet been optimized for application scenarios of unmanned retail. Currently, classification and object detection datasets do not exist that focus on unmanned retail solely. Therefore, in order to promote unmanned retail applications by using deep learning-based classification and object detection, in this article we collected more than 30 000 images of unmanned retail containers using a refrigerator affixed with different cameras under both static and dynamic recognition environments. These images were categorized into ten kinds of beverages. After manual labeling, images in our constructed dataset contained 155 153 instances, each of which was annotated with a bounding box. We performed extensive experiments on this dataset using ten state-of-the-art deep learning-based models. Experimental results indicate great potential of using these deep learning-based models for real-world smart unmanned vending machines. Haijun Zhang 0002, Donghai Li, Yuzhu Ji, Haibin Zhou, Weiwei Wu 0001, Kai Liu 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2019 | Deep Learning-based Beverage Recognition for Unmanned Vending Machines: An Empirical StudyabstractIn recent years, deep learning techniques have been commonly used in the fields of image processing and computer vision. With the popularity of deep learning models, researchers have developed many effective object detection methods. Unmanned retail applications start to utilizing object detection algorithms for changing traditional retail modes. Until now, there is no public datasets for object detection in unmanned retail application environments. Moreover, state-of-the-art deep learning-based object detection models have not yet been examined in this application scenario. In this paper, we compiled a large-scale dataset which contains over 30,000 images captured in a refrigerator equipped with different cameras. 10 kinds of beverages were utilized for targeted objects. An empirical study on this dataset is performed by using several recent developed deep learning models. Results demonstrate the effectiveness of using deep learning techniques real-life unmanned retail environments. Haijun Zhang 0002, Donghai Li, Yuzhu Ji, Haibin Zhou, Weiwei Wu 0001 |
INDIN | 3 |
| 2019 | Learning-based Object Detection in High Resolution UAV Images: An Empirical StudyabstractDeep learning-based methods are continuously boosting the performance of detecting objects in natural images. On the contrary, detecting objects in Unmanned Aerial Vehicle (UAV) images remains to be a difficult task in the field of computer vision, due to the challenge of training a well-performed detection model working on UAV images which usually contain instances with varied orientations, scales, and contours, etc. Furthermore, only a few researchers have focused on this field, probably because of difficulties in UAV data acquisition and labelling. Inspired by this, we collected a large-scale dataset with multi-scale and high-resolution UAV images, named MOHR, which contains 10,631 images captured by a UAV affixed with three kinds of cameras. Since these images were captured in a suburban environment, we manually annotated five classes of objects, including car, truck, building, collapse and flood damage. An empirical study was then conducted by adopting six advanced object detection methods all of which are based on deep learning technologies. The results indicate the great potential of these evaluated object detection models, but also reveal that the research on such a challenging UAV dataset using current deep learning techniques is far reaching. Haijun Zhang 0002, Mingshan Sun, Yuzhu Ji, Shichao Xu, Weihan Cao |
INDIN | 3 |
| 2019 | Graph model-based salient object detection using objectness and multiple saliency cues
Yuzhu Ji, Haijun Zhang 0002, Kuo-Kun Tseng, Tommy W. S. Chow, Q. M. Jonathan Wu |
Neurocomputing | 1 |
| 2019 | Toward AI fashion design: An Attribute-GAN model for clothing match
Haijun Zhang 0002, Yuzhu Ji, Q. M. Jonathan Wu |
Neurocomputing | 3 |
| 2019 | Sitcom-star-based clothing retrieval for video advertising: a deep learning framework
Haijun Zhang 0002, Yuzhu Ji, Wang Huang |
Neural Comput. Appl. | 2 |
| 2018 | Sitcom-Stars Oriented Video Advertising via Clothing Retrieval
Haijun Zhang 0002, Yuzhu Ji, Wang Huang |
DASFAA (2) | 2 |
| 2018 | Saliency detection via conditional adversarial image-to-image network
Yuzhu Ji, Haijun Zhang 0002, Q. M. Jonathan Wu |
Neurocomputing | 1 |
| 2018 | Salient object detection via multi-scale attention CNN
Yuzhu Ji, Haijun Zhang 0002, Q. M. Jonathan Wu |
Neurocomputing | 1 |
| 2017 | Understanding Subtitles by Character-Level Sequence-to-Sequence LearningabstractThis paper presents a character-level seque-nce-to-sequence learning method, RNNembed. This method allows the system to read raw characters, instead of words generated by preprocessing steps, into a pure single neural network model under an end-to-end framework. Specifically, we embed a recurrent neural network into an encoder–decoder framework and generate character-level sequence representation as input. The dimension of input feature space can be significantly reduced as well as avoiding the need to handle unknown or rare words in sequences. In the language model, we improve the basic structure of a gated recurrent unit by adding an output gate, which is used for filtering out unimportant information involved in the attention scheme of the alignment model. Our proposed method was examined in a large-scale dataset on an English-to-Chinese translation task. Experimental results demonstrate that the proposed approach achieves a translation performance comparable, or close, to conventional word-based and phrase-based systems. Haijun Zhang 0002, Yuzhu Ji, Heng Yue |
IEEE Trans. Ind. Informatics | 3 |
| 2016 | Distributed diffusion nonnegative LMS algorithm over sensor networksabstractSince most distributed estimation algorithms only try to achieve high estimation precision while ignoring the positive-negative problem of components in the true parameter, estimation using these methods may be physically absurd and uninterpretable. In order to avoid erroneous results, we need to add a nonnegative constraint on the parameter to be estimated. In this paper, we propose a novel distributed diffusion nonnegative LMS algorithm with regularization for estimating some specific parameter. The algorithm keeps the non-negativity of all components in the parameter in the adaptation process. Simulations results illustrate the advantage of our algorithm in the low steady MSD level and high convergence rate. Wei Huang 0015, Yuzhu Ji, Shengyong Chen |
INDIN | 3 |
| 2016 | Classifying vehicles with convolutional neural network and feature encodingabstractVehicle type recognition has many applications in video surveillance, urban traffic management and automatic driving. This paper presents a new vehicle type recognition method using feature encoding combined with Convolutional Neural Network (CNN). This method uses the CNN to learn the properties of the high-level image features. It is able to largely compensate the information loss if we use feature encoding solely. By contrast, to achieve satisfactory classification results, feature encoding algorithms do not need a large number of training samples. Thus, it can help CNN reduce the number of training samples. Therefore, we propose a hybrid algorithm by integrating method on vehicle type recognition in comparison to CNN, feature encoding algorithms and other competitive methods. Shuang Wang 0005, Zhengqi Li, Haijun Zhang 0002, Yuzhu Ji, Yan Li 0040 |
INDIN | 4 |
| 2016 | A character-level sequence-to-sequence method for subtitle learningabstractThis paper presents a character-level sequence-to-sequence learning method, RNNembed. Specifically, we embed a Recurrent Neural Network (RNN) into an encoder-decoder framework and generate character-level sequence representation as input. The dimension of input feature space can be significantly reduced as well as avoiding the need to handle unknown or rare words in sequences. In the language model, we improve the basic structure of a Gated Recurrent Unit (GRU) by adding an output gate, which is used for filtering out unimportant information involved in the attention scheme of the alignment model. Our proposed method was examined in a large-scale dataset on a task of English-to-Chinese translation. Experimental results demonstrate that the proposed approach achieves a translation performance comparable, or close, to conventional word-based and phrase-based systems. Haijun Zhang 0002, Yuzhu Ji, Heng Yue |
INDIN | 3 |
| 2016 | A Triple Wing Harmonium Model for Movie RecommendationabstractA new triple wing harmonium (TWH) model that integrates text metadata into a low-dimensional semantic space is proposed for the application of content-based movie recommendation. The text metadata considered here include movie synopsis, actor list, and user comments. We develop a new TWH model projecting these multiple textual features into low-dimensional latent topics with different probability distribution assumptions. A contrastive divergence (CD) algorithm is used for efficient learning and inference. Experimental results suggest that the proposed method performs better than the state-of-the-art algorithms for movie recommendation. Haijun Zhang 0002, Yuzhu Ji, Yunming Ye |
IEEE Trans. Ind. Informatics | 2 |
| 2015 | Content-based movie recommending using a Triple Wing Harmonium modelabstractA content-based movie recommender by using a Triple Wing Harmonium (TWH) model is proposed. TWH integrates text metadata into a low dimensional semantic space. movie synopsis, actor list and user comments are considered as the text metadata. A new TWH model is developed by projecting these multiple textual features into low dimensional latent topics. We have used a contrastive divergence algorithm for efficient learning and inference. Experimental results show that the proposed method performs better than the state-of-the-art content-based algorithms for movie recommendation. Haijun Zhang 0002, Yuzhu Ji, Yunming Ye |
INDIN | 3 |