EDBT 2026 Demo / reviewers in the wild / expert
Yijun Huang
dblp:89/2942
· DBLP profile ↗
14ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 6 since 2021Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Intelligent Master Station Abnormality Monitoring and Deep Diagnosis Framework Based on Residual Networks and Rule TemplatesabstractWith the rapid development of smart grids, the intelligent metering master station has become a critical control node whose operational stability directly affects data acquisition and power scheduling. However, due to the complexity of operating environments, the master station is prone to various types of abnormalities. Traditional rule-based detection methods struggle to handle diverse and implicit abnormal patterns, while deep learning methods, although powerful in modeling capabilities, often lack interpretability.To address these issues, this paper proposes a deep diagnosis framework named Residual and Rule Template-based Deep Diagnosis (RaRT-Diag), which integrates residual networks with expert-defined rule templates. The framework first uses a residual network to extract features from master station data and perform initial abnormality detection. Then, it employs rule templates derived from expert knowledge to refine the classification and conduct logical reasoning. This two-stage design realizes a collaborative mechanism of “coarse recognition[Formula: see text] [Formula: see text] [Formula: see text]fine diagnosis”, enhancing both detection accuracy and interpretability. Experiments conducted on real-world datasets demonstrate that RaRT-Diag outperforms traditional rule systems and standalone deep models in multi-type anomaly detection tasks, achieving higher accuracy and F1-score, and exhibiting strong generalization and practical value. Yijun Huang, Xingyuan Fan |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2026 | Flying Vehicle Detection Under Complex Conditions With RGB-Infrared Imagery: A Large-Scale Open-Source Suite and Benchmark ApproachabstractWhile coordination among multiple flying vehicles improves aerial logistics efficiency, safe and orderly operation requires robust detection and collision-avoidance capabilities. These requirements apply to passenger aircraft as well as to urban traffic and maritime environments, including emerging underwater flying vehicles. However, existing detection methods often fail under challenging conditions such as low illumination or cluttered backgrounds. Their progress is further constrained by the lack of large-scale benchmarks and the high computational and memory costs required to achieve high accuracy, which limits their deployment in resource-constrained scenarios, such as air-to-air collision avoidance in aerial vehicles. To address this gap, we introduce FT55k, an open-source benchmark comprising over 55,000 annotated RGB and infrared images across diverse environments. We further provide baseline approaches tailored for platforms with different computational demands. Extensive experiments on FT55k and three public datasets demonstrate the superior accuracy and efficiency of our methods compared with state-of-the-art approaches. Notably, our approach is the first flying vehicle detection method with a computational cost below 0.5 BFLOPs, achieving real-time performance at 62.3 FPS on an edge-computing device. This work presents the first comprehensive benchmark for flying vehicle detection in complex environments, establishing a practical and scalable foundation for future research and deployment in intelligent transportation safety. Our datasets is publicly accessible athttps://github.com/chriszxk/Flying-Vehicle-Detection Xunkuai Zhou, Yijun Huang, Li Li 0008, Jie Chen 0003, Ben M. Chen |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | SE-STDGNN: A Self-Evolving Spatial-Temporal Directed Graph Neural Network for Multi-Vehicle Trajectory PredictionabstractVehicle trajectory prediction (VTP) is essential for microscopic traffic risk assessment, autonomous vehicle navigation, and traffic behavior analysis. Related research leveraging learning-based methodologies has yielded notable success on various benchmark trajectory datasets. However, these models often experience performance degradation when faced with dynamic changes in traffic conditions such as vehicle density, road types, and weather conditions, as they have not been exposed to these variations during the training process. To effectively address the need for real-time adaptation in dynamic traffic scenarios, we propose a novel framework titled self-evolving spatial-temporal directed graph neural network (SE-STDGNN). This model utilizes evolving graph convolution networks (EvolveGCNs) to aggregate spatial-temporal features of vehicles and their neighbors, which are then utilized by a trajectory prediction module to forecast future trajectories. Further, a self-evolving mechanism is introduced to adjust model parameters dynamically in the real-time operation. The efficacy of SE-STDGNN is validated using the public vehicle trajectory dataset AD4CHE. Bingxin Han, Yijun Huang, Xi Chen 0104, Ben M. Chen |
ICRA | 3 |
| 2025 | Multi-View Stereo with Geometric Encoding for Dense Scene ReconstructionabstractMulti-view stereo (MVS) implicitly encodes photometric and geometric cues into the cost volume for multi-view correspondence matching, transferring insufficient geometric cues essential to depth estimation and reconstruction. This paper proposes GE-MVS, a novel multi-view stereo network with geometric encoding for more accurate and complete depth estimation and point cloud reconstruction. First, the cross-view adaptive cost volume aggregation module is proposed to strengthen multi-view geometric cues encoding during cost volume construction. Then, the depth consistency optimization is performed in the 3D point space during learning by invoking ground-truth depth cues from adjacent views. Finally, the surface normal geometries are explicitly encoded to refine the sampled depth hypotheses to be consistent in the local neighbor regions. Extensive experiments on the standard MVS benchmarks including DTU, Tanks and Temples, and BlendedMVS demonstrate the state-of-the-art depth estimation and point cloud reconstruction performance of GE-MVS. The GE-MVS is further deployed in real-world experiments for UAV-based large-scale reconstruction, where our method outperforms the prevalent industrial reconstruction solutions concerning reconstruction efficiency and efficacy. Our project page is: https://cuhk-usr-group.github.io/GE-MVS/ Guidong Yang, Junjie Wen 0001, Benyun Zhao, Qingxiang Li, Yijun Huang, Lei Lei 0010, Xi Chen 0104, Alan H. F. Lam, Ben M. Chen |
ICRA | 6 |
| 2025 | End-to-End Underwater Multi-View Stereo for Dense Scene ReconstructionabstractRecent advancements in learning-based multi-view stereo (MVS) have demonstrated significant improvements over traditional counterpart, primarily due to the extensive availability of multi-view training images with ground-truth metric depths in the terrestrial in-air domain. However, underwater multi-view stereo (UwMVS) faces substantial challenges arising from the domain gap between in-air and underwater environments, leading to degraded performance when applying in-air MVS models to underwater scenarios. Furthermore, the progress of learning-based UwMVS methods has been hindered by the scarcity of underwater multi-view images with ground-truth depth maps and point clouds. In this paper, we address these challenges by introducing a physically-guided approach for synthesizing underwater multi-view images and present the first large-scale UwMVS dataset for end-to-end training and evaluation of learning-based UwMVS methods. Furthermore, we propose a novel UwMVS network that enhances geometric cue encoding to achieve more accurate and complete point cloud reconstruction. Extensive experiments on our dataset and real-world underwater scenes demonstrate that our dataset enables the trained models for underwater dense reconstruction and that our method achieves state-of-the-art performance in underwater reconstruction. Dataset, code and appendix are available at: https://cuhk-usr-group.github.io/UwMVS/ Guidong Yang, Junjie Wen 0001, Benyun Zhao, Qingxiang Li, Yijun Huang, Lei Lei 0010, Xi Chen 0104, Alan H. F. Lam, Ben M. Chen |
ICRA | 5 |
| 2025 | FHGS: Feature-Homogenized Gaussian SplattingabstractScene understanding based on 3D Gaussian Splatting (3DGS) has recently achieved notable advances. Although 3DGS related methods have efficient rendering capabilities, they fail to address the inherent contradiction between the anisotropic color representation of gaussian primitives and the isotropic requirements of semantic features, leading to insufficient cross-view feature consistency.
To overcome the limitation, we proposes FHGS (Feature-Homogenized Gaussian Splatting), a novel 3D feature distillation framework inspired by physical models, which freezes and distills 2D pre-trained features into 3D representations while preserving the real-time rendering efficiency of 3DGS.
Specifically, our FHGS introduces the following innovations: Firstly, a universal feature fusion architecture is proposed, enabling robust embedding of large-scale pre-trained models' semantic features (e.g., SAM, CLIP) into sparse 3D structures.
Secondly, a non-differentiable feature fusion mechanism is introduced, which enables semantic features to exhibit viewpoint independent isotropic distributions. This fundamentally balances the anisotropic rendering of gaussian primitives and the isotropic expression of features; Thirdly, a dual-driven optimization strategy inspired by electric potential fields is proposed, which combines external supervision from semantic feature fields with internal primitive clustering guidance. This mechanism enables synergistic optimization of global semantic alignment and local structural consistency.
Extensive comparison experiments with other state-of-the-art methods on benchmark datasets demonstrate that our FHGS exhibits superior reconstruction performance in feature fusion, noise suppression, and geometric precision, while maintaining a significantly lower training time.
This work establishes a novel Gaussian Splatting data structure, offering practical advancements for real-time semantic mapping, 3D stylization, and Vision-Language Navigation (VLN).
Our code and additional results are available on our project page:https://fhgs.cuastro.org/. Qigeng Duan, Benyun Zhao, Mingqiao Han, Yijun Huang, Ben M. Chen |
NeurIPS | 4 |
| 2024 | Det-Recon-Reg: An Intelligent Framework Towards Automated Large-Scale Infrastructure InspectionabstractVisual inspection plays a predominant role in inspecting infrastructure surface. However, the generalization of existing visual inspection systems to large-scale real-world scenes remains challenging. In this paper, we introduce Det-Recon-Reg, an intelligent framework separating the complex inspection procedure into three stages: Detect, Reconstruct, and Register. (1) For defect detection (Detect), we present the first high-resolution defect dataset tailored for large-scale defect detection. Based on the dataset, we evaluate the most effective real-time object detection algorithms and push the boundary by proposing CUBIT-Net for real-world defect inspection. (2) For infrastructure reconstruction (Reconstruct), we propose a learning-based multi-view stereo (MVS) network to adapt to large-scale scenes, taking as input the multi-view images and outputting the point cloud reconstruction, where its performance has been validated on the standard MVS datasets, including BlendedMVS, DTU, and Tanks and Temples datasets. (3) For defect localization (Register), we propose an effective registration method based on the geographic information system that registers the detected defects onto the reconstructed infrastructure model to establish a global reference for maintenance measures. The real-world experiments further verify the effectiveness and efficiency of our proposed framework. More details about our proposed dataset, code, and appendix are available on our project page: https://cuhk-usr-group.github.io/large-scale-inspect-framework/. Guidong Yang, Jihan Zhang, Benyun Zhao, Chuanxiang Gao, Yijun Huang, Junjie Wen 0001, Qingxiang Li, Jerry Tang, Xi Chen 0104, Ben M. Chen |
IROS | 5 |
| 2024 | Exploring Structural Sparsity of Coil Images from 3-Dimensional Directional Tight Framelets for SENSE ReconstructionabstractAbstract. Each coil image in a parallel magnetic resonance imaging (pMRI) system is an imaging slice modulated by the corresponding coil sensitivity. These coil images, structurally similar to each other, are stacked together as 3-dimensional (3D) image data, and their sparsity property can be explored via 3D directional Haar tight framelets. The features of the 3D image data from the 3D framelet systems are utilized to regularize sensitivity encoding (SENSE) pMRI reconstruction. Accordingly, a so-called SENSE3d algorithm is proposed to reconstruct images of high quality from the sampled [Formula: see text]-space data with a high acceleration rate by decoupling effects of the desired image (slice) and sensitivity maps. Since both the imaging slice and sensitivity maps are unknown, this algorithm repeatedly performs a slice step followed by a sensitivity step by using updated estimations of the desired image and the sensitivity maps. In the slice step, for the given sensitivity maps, the estimation of the desired image is viewed as the solution to a convex optimization problem regularized by the sparsity of its 3D framelet coefficients of coil images. This optimization problem, involving data from the complex field, is solved by a primal-dual three-operator splitting (PD3O) method. In the sensitivity step, the estimation of sensitivity maps is modeled as the solution to a Tikhonov-type optimization problem that favors the smoothness of the sensitivity maps. This corresponding problem is nonconvex and could be solved by a forward-backward splitting method. Experiments on real phantoms and in vivo data show that the proposed SENSE3d algorithm can explore the sparsity property of the imaging slices and efficiently produce reconstructed images of high quality with reduced aliasing artifacts caused by high acceleration rate, additive noise, and the inaccurate estimation of each coil sensitivity. To provide a comprehensive picture of the overall performance of our SENSE3d model, we provide the quantitative index (HaarPSI) and comparisons to some deep learning methods such as VarNet and fastMRI-UNet. Yanran Li, Raymond Chan 0001, Lixin Shen, Xiaosheng Zhuang, Risheng Wu, Yijun Huang |
SIAM J. Imaging Sci. | 6 |
| 2018 | New Balanced Active Learning Model and Optimization AlgorithmabstractIt is common in machine learning applications that unlabeled data are abundant while acquiring labels is extremely difficult. In order to reduce the cost of training model while maintaining the model quality, active learning provides a feasible solution. Instead of acquiring labels for random samples, active learning methods carefully select the data to be labeled so as to alleviate the impact from the redundancy or noise in the selected data and improve the trained model performance. In early stage experimental design, previous active learning methods adopted data reconstruction framework, such that the selected data maintained high representative power. However, these models did not consider the data class structure, thus the selected samples could be predominated by the samples from major classes. Such mechanism fails to include samples from the minor classes thus tends to be less "representative". To solve this challenging problem, we propose a novel active learning model for the early stage of experimental design. We use exclusive sparsity norm to enforce the selected samples to be (roughly) evenly distributed among different groups. We provide a new efficient optimization algorithm and theoretically prove the optimal convergence rate O(1/{T^2}). With a simple substitution, we reduce the computational load of each iteration from O(n^3) to O(n^2), which makes our algorithm more scalable than previous frameworks. Xiaoqian Wang 0001, Yijun Huang, Heng Huang 0001 |
IJCAI | 2 |
| 2018 | Exclusive Sparsity Norm Minimization With Random Groups via Cone ProjectionabstractMany practical applications such as gene expression analysis, multitask learning, image recognition, signal processing, and medical data analysis pursue a sparse solution for the feature selection purpose and particularly favor the nonzeros evenly distributed in different groups. The exclusive sparsity norm has been widely used to serve to this purpose. However, it still lacks systematical studies for exclusive sparsity norm optimization. This paper offers two main contributions from the optimization perspective: 1) we provide several efficient algorithms to solve exclusive sparsity norm minimization with either smooth loss or hinge loss (nonsmooth loss). All algorithms achieve the optimal convergence rate . ( is the iteration number.) To the best of our knowledge, this is the first time to guarantee such convergence rate for the general exclusive sparsity norm minimization and 2) when the group information is unavailable to define the exclusive sparsity norm, we propose to use the random grouping scheme to construct groups and prove that if the number of groups is appropriately chosen, the nonzeros (true features) would be grouped in the ideal way with high probability. Empirical studies validate the efficiency of the proposed algorithms, and the effectiveness of random grouping scheme on the proposed exclusive support vector machine formulation. Yijun Huang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | CHI: A contemporaneous health index for degenerative disease monitoring using longitudinal measurements
Yijun Huang, Heather L. Evans, William B. Lober, Yu Cheng 0001, Xiaoning Qian, Ji Liu 0002, Shuai Huang 0001 |
J. Biomed. Informatics | 1 |
| 2016 | On Benefits of Selection Diversity via Bilevel Exclusive SparsityabstractSparse feature (dictionary) selection is critical for various tasks in computer vision, machine learning, and pattern recognition to avoid overfitting. While extensive research efforts have been conducted on feature selection using sparsity and group sparsity, we note that there has been a lack of development on applications where there is a particular preference on diversity. That is, the selected features are expected to come from different groups or categories. This diversity preference is motivated from many real-world applications such as advertisement recommendation, privacy image classification, and design of survey. In this paper, we proposed a general bilevel exclusive sparsity formulation to pursue the diversity by restricting the overall sparsity and the sparsity in each group. To solve the proposed formulation that is NP hard in general, a heuristic procedure is proposed. The main contributions in this paper include: 1) A linear convergence rate is established for the proposed algorithm, 2) The provided theoretical error bound improves the approaches such as L1norm and L0types methods which only use the overall sparsity and the quantitative benefits of using the diversity sparsity is provided. To the best of our knowledge, this is the first work to show the theoretical benefits of using the diversity sparsity, 3) Extensive empirical studies are provided to validate the proposed formulation, algorithm, and theory. Haichuan Yang, Yijun Huang, Lam Tran, Ji Liu 0002, Shuai Huang 0001 |
CVPR | 2 |
| 2016 | A Comprehensive Linear Speedup Analysis for Asynchronous Stochastic Parallel Optimization from Zeroth-Order to First-OrderabstractAsynchronous parallel optimization received substantial successes and extensive attention recently. One of core theoretical questions is how much speedup (or benefit) the asynchronous parallelization can bring to us. This paper provides a comprehensive and generic analysis to study the speedup property for a broad range of asynchronous parallel stochastic algorithms from the zeroth order to the first order methods. Our result recovers or improves existing analysis on special cases, provides more insights for understanding the asynchronous parallel behaviors, and suggests a novel asynchronous parallel zeroth order method for the first time. Our experiments provide novel applications of the proposed asynchronous parallel zeroth order method on hyper parameter tuning and model blending problems. Xiangru Lian, Huan Zhang 0001, Cho-Jui Hsieh, Yijun Huang, Ji Liu 0002 |
NIPS | 4 |
| 2015 | Asynchronous Parallel Stochastic Gradient for Nonconvex OptimizationabstractThe asynchronous parallel implementations of stochastic gradient (SG) have been broadly used in solving deep neural network and received many successes in practice recently. However, existing theories cannot explain their convergence and speedup properties, mainly due to the nonconvexity of most deep learning formulations and the asynchronous parallel mechanism. To fill the gaps in theory and provide theoretical supports, this paper studies two asynchronous parallel implementations of SG: one is on the computer network and the other is on the shared memory system. We establish an ergodic convergence rate $O(1/\sqrt{K})$ for both algorithms and prove that the linear speedup is achievable if the number of workers is bounded by $\sqrt{K}$ ($K$ is the total number of iterations). Our results generalize and improve existing analysis for convex minimization. Xiangru Lian, Yijun Huang, Yuncheng Li, Ji Liu 0002 |
NIPS | 2 |