EDBT 2026 Demo / reviewers in the wild / expert
Jinhui Yuan
dblp:58/3397
· DBLP profile ↗
21ranked-venue papers
9as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 first-authorArtificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorSecurity and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Comprehensive Deadlock Prevention for GPU Collective CommunicationabstractDistributed deep neural network training necessitates efficient GPU collective communications, which are inherently susceptible to deadlocks. GPU collective deadlocks arise easily in distributed deep learning applications when multiple collectives circularly wait for each other. GPU collective deadlocks pose a significant challenge to the correct functioning and efficiency of distributed deep learning, and no general effective solutions are currently available. Only in specific scenarios, ad-hoc methods, making an application invoke collectives in a consistent order across GPUs, can be used to prevent circular collective dependency and deadlocks. Lichen Pan, Yongquan Fu, Jinhui Yuan, Rongkai Zhang 0005, Pengze Li |
EuroSys | 4 |
| 2025 | AGHINT: Attribute-guided representation learning on heterogeneous information networks with transformer
Jinhui Yuan, Shan Lu 0014, Peibo Duan, Jieyue He |
Knowl. Based Syst. | 1 |
| 2025 | Research on Desertification Monitoring and Vegetation Refinement Extraction Methods Based on the Synergy of Multisource Remote Sensing ImageryabstractDue to over-exploitation by humans and global climate change, desertification has become an increasingly severe issue, seriously threatening the stability of ecosystems and the sustainable development of resources. Therefore, this study focuses on the Hangjin Banner region in Inner Mongolia, using satellite remote sensing and remote aerial vehicles (RAV) remote sensing technology. Through wide-area coverage, long-term monitoring, multiscale analysis, and high-precision interpretation, the study demonstrates the strong synergistic effects of “multiscale interpretation” and “data fusion applications,” systematically carrying out desertification monitoring grading and refined vegetation extraction. First, to address the problem that the information dimension of a single index is insufficient and it is difficult to reflect the development trend of desertification, the normalized difference vegetation index (NDVI)-albedo feature space applicable to the desert environment is inversely performed based on Landsat 8 satellite images from 2009 to 2023. Then, on the basis of the feature space, the desertification difference index (DDI), which realizes the wide-area desertification monitoring grading and spatio-temporal evolution analysis of the study area, and the hue-saturation-lightness greenway enhanced vegetation index (HSLGEVI), which has stronger applicability and stability in desert environments, were constructed based on the HSL color space and the hue tuning algorithm. This index can effectively overcome the limitations of the RGB vegetation index, clearly delineate the canopy edge of desert vegetation, and accurately extract surface meadow vegetation with lower chlorophyll content. To test the effectiveness of the HSLGEVI, the widely used and validated excess green index (EXG), vegetation difference vegetation index (VDVI), modified green-red vegetation index (MGRVI), and red-green–blue vegetation index (RGBVI) were selected for comparison. The results show that the accuracy of HSLGEVI is better than that of other indices, with overall accuracy and${F}1$-score remaining above 90%. It reduces the impact of the RGB color space vegetation index on the accuracy of vegetation extraction, effectively overcoming misclassification and omission issues, and providing a reliable monitoring mechanism for desertification control in the Hangjin Banner area. Zhenqi Song, Yuefeng Lu, Jinhui Yuan, Miao Lu, Dengkuo Sun |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | GASF-ConvNeXt-TF Algorithm for Perimeter Security Disturbance Identification Based on Distributed Optical Fiber Sensing Systemabstractφ-OTDR technology can transform the fiber optic cable into a large-scale sensor array for distributed acoustic sensing (DAS), which is an emerging infrastructure for the Internet of Things. However, it’s limitated in event recognition capability, which is a major factor preventing its practical field application. This paper proposes a perturbation recognition algorithm based on GASF-ConvNeXt-TF with fast process and high recognition accuracy. Firstly, GASF (Gramian angular summation field) algorithm is used to encode external disturbance signal to transform the one-dimensional time series signal into a more concentrated two-dimensional image feature. Then the CNN model ConvNeXttiny network is applied as the classifier. In order to prevent the weight gradient from oscillating back and forth during network training process, a cosine annealing algorithm is introduced to control the decay of the learning rate. Meanwhile, transfer learning is used to further optimize the network model, resulting in higher classification accuracy and faster convergence. Finally, two different experimental scenarios are arranged in a total length of 2.2 kilometers of optical fiber cable, and six different disturbance events (shaking, kicking, knocking, trampling, wheel rolling, and impacting) are set. Different from previous perimeter security disturbance identification experiments, not only single-point disturbance recognition is performed, but also two points disturbances are simultaneously recognized, and all have good overall identification accuracy. The overall recognition accuracy of the six disturbance events in single and multiple points experiments are 99.3% and 98.3%, respectively, with an average recognition time of 0.103s. The proposed technique has potential application in infrastructures structure health monitoring, such as factories, airports, energy pipeline and highway. Ya-Jun Wang, Wen Zhuo, Bin Liu 0048, Juan Liu 0010, Xingdao He, Jinhui Yuan, Qiang Wu 0005 |
IEEE Internet Things J. | 9 |
| 2024 | AutoDDL: Automatic Distributed Deep Learning With Near-Optimal Bandwidth CostabstractRecent advances in deep learning are driven by the growing scale of computation, data, and models. However, efficiently training large-scale models on distributed systems requires an intricate combination of data, operator, and pipeline parallelism, which exerts heavy burden on machine learning practitioners. To this end, we propose AutoDDL, a distributed training framework that automatically explores and exploits new parallelization schemes with near-optimal bandwidth cost. AutoDDL facilitates the description and implementation of different schemes by utilizing OneFlow'sSplit,Broadcast, andPartial Sum(SBP) abstraction. AutoDDL is equipped with an analytical performance model combined with a customized Coordinate Descent algorithm, which significantly reduces the scheme searching overhead. We conduct evaluations on Multi-Node-Single-GPU and Multi-Node-Multi-GPU machines using different models, including VGG and Transformer. Compared to the expert-optimized implementations, AutoDDL reduces the end-to-end training time by up to 31.1% and 10% for Transformer and up to 17.7% and 71.5% for VGG on the two parallel systems, respectively. Jinfan Chen, Shigang Li 0002, Jinhui Yuan, Torsten Hoefler |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2024 | Swift: Expedited Failure Recovery for Large-Scale DNN TrainingabstractAs the size of deep learning models gets larger and larger, training takes longer time and more resources, making fault tolerance more and more critical. Existing state-of-the-art methods like CheckFreq and Elastic Horovod need to back up a copy of the model state (i.e., parameters and optimizer states) in memory, which is costly for large models and leads to non-trivial overhead. This article presentsSwift, a novel recovery design for distributed deep neural network training that significantly reduces the failure recovery overhead without affecting training throughput and model accuracy. Instead of making an additional copy of the model state,Swiftresolves the inconsistencies of the model state caused by the failure and exploits the replicas of the model state in data parallelism for failure recovery. We propose a logging-based approach when replicas are unavailable, which records intermediate data and replays the computation to recover the lost state upon a failure. The re-computation is distributed across multiple machines to accelerate failure recovery further. We also log intermediate data selectively, exploring the trade-off between recovery time and intermediate data storage overhead. Evaluations show thatSwiftsignificantly reduces the failure recovery time and achieves similar or better training throughput during failure-free execution compared to state-of-the-art methods without degrading final model accuracy.Swiftcan also achieve up to 1.16x speedup in total training time compared to state-of-the-art methods. Yuchen Zhong, Guangming Sheng, Jinhui Yuan, Chuan Wu 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2023 | Coop: Memory is not a CommodityabstractTensor rematerialization allows the training of deep neural networks (DNNs) under limited memory budgets by checkpointing the models and recomputing the evicted tensors as needed. However, the existing tensor rematerialization techniques overlook the memory system in deep learning frameworks and implicitly assume that free memory blocks at different addresses are identical. Under this flawed assumption, discontiguous tensors are evicted, among which some are not used to allocate the new tensor. This leads to severe memory fragmentation and increases the cost of potential rematerializations.
To address this issue, we propose to evict tensors within a sliding window to ensure all evictions are contiguous and are immediately used. Furthermore, we proposed cheap tensor partitioning and recomputable in-place to further reduce the rematerialization cost by optimizing the tensor allocation.
We named our method Coop as it is a co-optimization of tensor allocation and tensor rematerialization. We evaluated Coop on eight representative DNNs. The experimental results demonstrate that Coop achieves up to $2\times$ memory saving and hugely reduces compute overhead, search latency, and memory fragmentation compared to the state-of-the-art baselines. Shihan Ma, Peihong Liu, Jinhui Yuan |
NeurIPS | 4 |
| 2023 | Swift: Expedited Failure Recovery for Large-Scale DNN TrainingabstractAs the size of deep learning models gets larger and larger, training takes longer time and more resources, making fault tolerance critical. Existing state-of-the-art methods like Check-Freq and Elastic Horovod need to back up a copy of the model state in memory, which is costly for large models and leads to non-trivial overhead. This paper presents Swift, a novel failure recovery design for distributed deep neural network training that significantly reduces the failure recovery overhead without affecting training throughput and model accuracy. Instead of making an additional copy of the model state, Swift resolves the inconsistencies of the model state caused by the failure and exploits replicas of the model state in data parallelism for failure recovery. We propose a logging-based approach when replicas are unavailable, which records intermediate data and replays the computation to recover the lost state upon a failure. Evaluations show that Swift significantly reduces the failure recovery time and achieves similar or better training throughput during failure-free execution compared to state-of-the-art methods without degrading final model accuracy. Yuchen Zhong, Guangming Sheng, Jinhui Yuan, Chuan Wu 0001 |
PPoPP | 4 |
| 2023 | Machine-Learning-Based Human Motion Recognition via Wearable Plastic-Fiber Sensing SystemabstractWearable human–machine interface (HMI) is a medium for information transmission and exchange between people and computers. It is widely used in the fields of human motion capture and recognition and augmented/virtual reality (AR/VR). This research proposes a wearable plastic-optical-fiber (POF) sensing system based on machine learning for human motion recognition. The wearable sports sleeve is designed and worn on the elbow and knee joints of human body. The wearable sensor system uses a D-shaped POF (DPOF) sensor, whose coefficient of determination (R 2) is 0.96496 and sensitivity is -0.7859% per degree. Support vector machines (SVMs), MobileNetV2 network, and transfer learning were used to identify six types of movement: walking, running, going upstairs, going downstairs, high leg lifts, and rope skipping. The accuracy of classification based on the four joint position monitoring can reach 98.28%, 98.94%, and 99.74%, respectively. The proposed POF wearable system has good applications for human motion state recognition and possesses great application potential in AR/VR. Bin Liu 0048, Yu-Lin Wang, Juan Liu 0010, Xingdao He, Jinhui Yuan, Qiang Wu 0005 |
IEEE Internet Things J. | 7 |
| 2021 | Theoretical analysis of PAM-N and M-QAM BER computation with single-sideband signal
Dongxu Lu, Xian Zhou 0001, Yuqiang Yang, Jiahao Huo, Jinhui Yuan, Keping Long, Changyuan Yu, Alan Pak Tao Lau, Chao Lu 0001 |
Sci. China Inf. Sci. | 5 |
| 2020 | Theoretical and numerical analyses for PDM-IM signals using Stokes vector receivers
Jiahao Huo, Xian Zhou 0001, Wei Huangfu, Jinhui Yuan, Huansheng Ning, Keping Long, Changyuan Yu, Alan Pak Tao Lau, Chao Lu 0001 |
Sci. China Inf. Sci. | 5 |
| 2015 | LightLDA: Big Topic Models on Modest Computer ClustersabstractWhen building large-scale machine learning (ML) programs, such as massive topic models or deep neural networks with up to trillions of parameters and training examples, one usually assumes that such massive tasks can only be attempted with industrial-sized clusters with thousands of nodes, which are out of reach for most practitioners and academic researchers. We consider this challenge in the context of topic modeling on web-scale corpora, and show that with a modest cluster of as few as 8 machines, we can train a topic model with 1 million topics and a 1-million-word vocabulary (for a total of 1 trillion parameters), on a document collection with 200 billion tokens --- a scale not yet reported even with thousands of machines. Our major contributions include: 1) a new, highly-efficient O(1) Metropolis-Hastings sampling algorithm, whose running cost is (surprisingly) agnostic of model size, and empirically converges nearly an order of magnitude more quickly than current state-of-the-art Gibbs samplers; 2) a model-scheduling scheme to handle the big model challenge, where each worker machine schedules the fetch/use of sub-models as needed, resulting in a frugal use of limited memory capacity and network bandwidth; 3) a differential data-structure for model storage, which uses separate data structures for high- and low-frequency words to allow extremely large models to fit in memory, while maintaining high inference speed. These contributions are built on top of the Petuum open-source distributed ML framework, and we provide experimental evidence showing how this development puts massive data and models within reach on a small cluster, while still enjoying proportional time cost reductions with increasing cluster size. Jinhui Yuan, Fei Gao 0018, Qirong Ho, Wei Dai 0003, Jinliang Wei, Xun Zheng, Eric P. Xing, Tie-Yan Liu, Wei-Ying Ma |
WWW | 1 |
| 2014 | HDROP: Detecting ROP Attacks Using Performance Monitoring Counters
Wenchang Shi, Jinhui Yuan, Bin Liang 0002 |
ISPEC | 4 |
| 2008 | Classifying What-Type Questions by Head Noun Tagging
Fangtao Li, Xian Zhang 0006, Jinhui Yuan, Xiaoyan Zhu 0001 |
COLING | 3 |
| 2008 | Scene understanding with discriminative structured predictionabstractSpatial priors play crucial roles in many high-level vision tasks, e.g. scene understanding. Usually, learning spatial priors relies on training a structured output model. In this paper, two special cases of discriminative structured output model, i.e. conditional random fields (CRFs) and max-margin Markov networks (M3N), are demonstrated to perform image scene understanding. The two models are empirically compared in a fair manner, i.e. using the common feature representation and the same optimization algorithm. Particularly, we adopt online exponentiated gradient (EG) algorithm to solve the convex duals of both models. We describe the general procedure of EG algorithm and present a two-stage training procedure to overcome the degeneration of EG when exact inference is intractable. Experiments on a large scale image region annotation task are carried out. The results show that both models yield encouraging results but CRFs slightly outperforms M3N. Jinhui Yuan, Jianmin Li 0001, Bo Zhang 0010 |
CVPR | 1 |
| 2007 | Gradual transition detection with conditional random fieldsabstractIn this paper, we view gradual transition detection as a sequence labeling problem and propose to use Conditional Random Fields (CRFs) for this purpose. CRFs is a state-of-the-art sequence labeling approach. It provides a unified way to integrate various useful clues to form a decision system. Moreover, it has principled way for parameter estimation and inference. Compared to rule-based approaches, gradual transition detection with CRFs requires fewer human interactions while designing the system. The experiments on TRECVID platform show that CRFs can achieve comparable performance to that of the state-of-the-art approaches. Jinhui Yuan, Jianmin Li 0001, Bo Zhang 0010 |
ACM Multimedia | 1 |
| 2007 | Exploiting spatial context constraints for automatic image region annotationabstractIn this paper we conduct a relatively complete study on how to exploit spatial context constraints for automated image region annotation. We present a straight forward method to regularize the segmented regions into 2D lattice layout, so that simple grid-structure graphical models can be employed to characterize the spatial dependencies. We show how to represent the spatial context constraints in various graphical models and also present the related learning and inference algorithms. Different from most of the existing work, we specifically investigate how to combine the classification performance of discriminative learning and the representation capability of graphical models. To reliably evaluate the proposed approaches, we create a moderate scale image set with region-level ground truth. The experimental results show that (i) spatial context constraints indeed help for accurate region annotation, (ii) the approaches combining the merits of discriminative learning and context constraints perform best, (iii) image retrieval can benefit from accurate region-level annotation. Jinhui Yuan, Jianmin Li 0001, Bo Zhang 0010 |
ACM Multimedia | 1 |
| 2007 | A Formal Study of Shot Boundary DetectionabstractThis paper conducts a formal study of the shot boundary detection problem. First, a general formal framework of shot boundary detection techniques is proposed. Three critical techniques, i.e., the representation of visual content, the construction of continuity signal and the classification of continuity values, are identified and formulated in the perspective of pattern recognition. Meanwhile, the major challenges to the framework are identified. Second, a comprehensive review of the existing approaches is conducted. The representative approaches are categorized and compared according to their roles in the formal framework. Based on the comparison of the existing approaches, optimal criteria for each module of the framework are discussed, which will provide practical guide for developing novel methods. Third, with all the above issues considered, we present a unified shot boundary detection system based on graph partition model. Extensive experiments are carried out on the platform of TRECVID. The experiments not only verify the optimal criteria discussed above, but also show that the proposed approach is among the best in the evaluation of TRECVID 2005. Finally, we conclude the paper and present some further discussions on what shot boundary detection can learn from other related fields Jinhui Yuan, Wujie Zheng, Jianmin Li 0001, Fuzong Lin, Bo Zhang 0010 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2006 | Learning concepts from large scale imbalanced data sets using support cluster machinesabstractThis paper considers the problem of using Support Vector Machines (SVMs) to learn concepts from large scale imbalanced data sets. The objective of this paper is twofold. Firstly, we investigate the effects of large scale and imbalance on SVMs. We highlight the role of linear non-separability in this problem. Secondly, we develop a both practical and theoretical guaranteed meta-algorithm to handle the trouble of scale and imbalance. The approach is named Support Cluster Machines (SCMs). It incorporates the informative and the representative under-sampling mechanisms to speedup the training procedure. The SCMs differs from the previous similar ideas in two ways, (a) the theoretical foundation has been provided, and (b) the clustering is performed in the feature space rather than in the input space. The theoretical analysis not only provides justification, but also guides the technical choices of the proposed approach. Finally, experiments on both the synthetic and the TRECVID data are carried out. The results support the previous analysis and show that the SCMs are efficient and effective while dealing with large scale imbalanced data sets. Jinhui Yuan, Jianmin Li 0001, Bo Zhang 0010 |
ACM Multimedia | 1 |
| 2005 | A unified shot boundary detection framework based on graph partition modelabstractIn this paper, we propose a unified shot boundary detection framework by extending the previous work of graph partition model with temporal constraints. To detect both the abrupt transitions (CUTs) and gradual transitions (GTs, excluding fade out/in) in a unified way, we incorporate temporal multi-resolution analysis into the model. Furthermore, instead of ad-hoc thresholding scheme, we construct a novel kind of feature to characterize shot transitions and employ support vector machine (SVM) with active leaning strategy to classify boundaries and non-boundaries. Extensive experiments have been carried out on the platform of TRECVID benchmark. The experimental results show that the proposed framework outperforms some others and achieves satisfactory results. Jinhui Yuan, Jianmin Li 0001, Fuzong Lin, Bo Zhang 0010 |
ACM Multimedia | 1 |
| 2005 | Graph Partition Model for Robust Temporal Data Segmentation
Jinhui Yuan, Bo Zhang 0010, Fuzong Lin |
PAKDD | 1 |