VLDB 2026 Research / reviewers in the wild / expert
He Guan
dblp:206/7476
· DBLP profile ↗
12ranked-venue papers
4as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Electronic design automation · 61% Integrated circuit design · 30% Energy-efficient computing · 9% | |
| Artificial intelligence
2 papers |
Generative modeling · 67% Vision and language · 33% | |
| Computer graphics and multimedia
2 papers |
Multimedia analysis and retrieval · 100% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Integrated circuit design
analog and mixed-signal circuits |
1.0 | 1 | 2026 | Multiscenario Coupled Model of GaN-on-Si HEMTs Epitaxial Design and Performance Improvement Based on Multiobjective Optimization Method · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2026 |
Electronic design automation
design optimization |
1.0 | 1 | 2026 | Multiscenario Coupled Model of GaN-on-Si HEMTs Epitaxial Design and Performance Improvement Based on Multiobjective Optimization Method · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2026 |
Electronic design automation
multi-objective optimization |
1.0 | 1 | 2026 | Multiscenario Coupled Model of GaN-on-Si HEMTs Epitaxial Design and Performance Improvement Based on Multiobjective Optimization Method · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2026 |
Multimedia analysis and retrieval
multimodal fusion |
0.4 | 2 | 2018 | Integrating Both Visual and Audio Cues for Enhanced Video Caption · AAAI 2018 CMCGAN: A Uniform Framework for Cross-Modal Visual-Audio Mutual Generation · AAAI 2018 |
Machine learning › Generative modeling
cross-modal generation |
0.3 | 1 | 2018 | CMCGAN: A Uniform Framework for Cross-Modal Visual-Audio Mutual Generation · AAAI 2018 |
Machine learning › Generative modeling
generative adversarial network |
0.3 | 1 | 2018 | CMCGAN: A Uniform Framework for Cross-Modal Visual-Audio Mutual Generation · AAAI 2018 |
Computer vision › Vision and language
video captioning |
0.3 | 1 | 2018 | Integrating Both Visual and Audio Cues for Enhanced Video Caption · AAAI 2018 |
Multimedia analysis and retrieval › multimodal fusion
audio-visual fusion |
0.3 | 1 | 2018 | Integrating Both Visual and Audio Cues for Enhanced Video Caption · AAAI 2018 |
Energy-efficient computing
thermal management |
0.3 | 1 | 2026 | Multiscenario Coupled Model of GaN-on-Si HEMTs Epitaxial Design and Performance Improvement Based on Multiobjective Optimization Method · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2026 |
Multimedia analysis and retrieval › multimodal learning
modality missing |
0.2 | 2 | 2018 | Integrating Both Visual and Audio Cues for Enhanced Video Caption · AAAI 2018 CMCGAN: A Uniform Framework for Cross-Modal Visual-Audio Mutual Generation · AAAI 2018 |
Methods — techniques the papers use, named apart from their topics
townsend multilayer film curvature theory · 1.0multi-objective optimization · 1.0weight sharing · 0.7multimodal memory · 0.7multimodal deep fusion · 0.7latent gaussian vector · 0.7adversarial loss · 0.7CycleGAN · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multiscenario Coupled Model of GaN-on-Si HEMTs Epitaxial Design and Performance Improvement Based on Multiobjective Optimization MethodabstractGallium Nitride High Electron Mobility Transistors (GaN-on-Si HEMTs) on silicon substrates exhibit significant potential in the field of radio-frequency devices due to its high frequency, high-speed, and low-cost advantages. Nevertheless, the thermal stress caused by the mismatch in lattice parameters and the coefficient of thermal expansion between the Si substrate and GaN epitaxial layer constrains the device performance. Due to the different temperature distributions in different scenarios, there is an urgent need for a model that can balance the contradiction between epitaxial growth and device operation for the joint design of GaN-on-Si HEMTs epitaxial structures. This work presents thermal stress formulas specifically designed for GaN epitaxial growth and device operation, based on Townsend’s theory of multilayer film curvature and stress. Furthermore, our work innovatively proposes a multi-scenario coupled GaN-on-Si HEMTs epitaxial structure design model based on multi-objective optimization. The model represents a combined objective function incorporating thermal stress experienced during device operation and epitaxial growth, as well as the total thermal resistance and bulk thermal conductivity of the epitaxial layer. Through simulation validation, the model has been proved to effectively calculate the optimized epitaxial structure parameters across multiple scenarios to meet diverse requirements. And it can effectively reduce the operating temperature. This model can accelerate the process of GaN epitaxial design and provide an effective method for enhancing the performance of GaN-on-Si HEMTs. Borui Deng, Yulong Fang, He Guan, Ziqiang Zeng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Failure Mode and Effects Analysis Method on the Air System of an Aircraft Turbofan Engine in Multi-Criteria Open Group Decision-Making EnvironmentabstractFailure mode and effects analysis (FMEA), an proactive risk management approach, has been widely applied in a variety of industries, especially in aircraft industry. In the process of implementation, the influence and uncertainty among different experts is inevitable. In order to handle the uncertainty in the assessments of FMEA experts, Dempster-Shafer evidence theory was introduced to FMEA for its flexibility and superiority in coping with uncertain and subjective assessments. However, traditional Dempster combination rule have difficulty in dealing with highly conflicting evidence that given from FMEA experts’ assessments. Moreover, experts themselves may influence each other though process such as chatting, judging, decision-making and voting. In this paper, we explore the problem of conflict evidence fusion from a correlation perspective among FMEA experts. We use ambiguity measure and Gaussian distribution to deal with the highly conflicting evidence. We use ambiguity measure to calculate the variance of Gaussian distribution. Then, we use Gaussian model to generalize expert assessments. After that, we use Dempster combination rule to fuze assessments from different experts. Finally, we calculate the risk priority number to rank the risk level of the FMEA items. The experiment results in the air system of an aircraft turbofan engine shows the efficiency and accuracy of the proposed method. Zixi Fei, Bingying Zhao, He Guan |
Cybern. Syst. | 6 |
| 2025 | VAG: A Uniform Model for Cross-Modal Visual-Audio Mutual GenerationabstractConsidering both audio and visual modalities is helpful for understanding a video. In the face of harsh environmental interference or signal packet loss, automatically compensating for audio and vision is a challenging task. We propose a dynamic cross-modal visual-audio mutual generation model (VAMG), which includes audio to visual conversion, visual to audio conversion, audio self-generation, and visual self-generation. VAMG jointly optimizes modal reconstruction and adversarial constraints, effectively solving the problems of structural alignment and signal compensation in incomplete videos. We conducted an instrument-oriented and pose-oriented cross-modal audio-visual mutual generation experiment on the sub-University of Rochester Musical Performance dataset to verify the effectiveness of the model. He Guan, Zhaoxiang Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Broadband High-Efficiency Continuous Class-F/F-1 Power Amplifiers with Active Second Harmonic Injection TechniqueabstractIn this paper, focusing on the power amplifier (PA) in ultra low power RF circuits and systems, a novel circuit topology with active second harmonic injection technique is derived and analyzed for the wideband high-efficiency continuous class-F/F-1(CCF/F-1) modes for the first time. By introducing an auxiliary PA operated at the second harmonic frequency of the desired band, the optimal phase shift parameter between the drain voltage and current of the CCF/F-1PA can be selected properly, leading to the improvement of the total efficiency and output power. Calculated results show that, with the proposed structure, the maximum output powers are 1.323 and 1.319 times larger than those in the traditional CCF and CCF-1modes, and the maximum drain efficiencies are increased to 98.16% and 93.16%, respectively. Finally, in order to verify the validity of the proposed circuit, a frequency-domain simulation has been presented. It is shown that the difference of intrinsic voltage and current waveforms, drain efficiency as well as output power between the results of theory and simulation is very small. Chang Liu 0044, He Guan, Hao Zhang 0076, Fadhel M. Ghannouchi |
ISCAS | 4 |
| 2024 | GRAMO: geometric resampling augmentation for monocular 3D object detectionabstractAbstract Data augmentation is widely recognized as an effective means of bolstering model robustness. However, when applied to monocular 3D object detection, non-geometric image augmentation neglects the critical link between the image and physical space, resulting in the semantic collapse of the extended scene. To address this issue, we propose two geometric-level data augmentation operators named Geometric-Copy-Paste (Geo-CP) and Geometric-Crop-Shrink (Geo-CS). Both operators introduce geometric consistency based on the principle of perspective projection, complementing the options available for data augmentation in monocular 3D. Specifically, Geo-CP replicates local patches by reordering object depths to mitigate perspective occlusion conflicts, and Geo-CS re-crops local patches for simultaneous scaling of distance and scale to unify appearance and annotation. These operations ameliorate the problem of class imbalance in the monocular paradigm by increasing the quantity and distribution of geometrically consistent samples. Experiments demonstrate that our geometric-level augmentation operators effectively improve robustness and performance in the KITTI and Waymo monocular 3D detection benchmarks. He Guan, Chunfeng Song, Zhaoxiang Zhang 0001 |
Frontiers Comput. Sci. | 1 |
| 2022 | Toward few-shot domain adaptation with perturbation-invariant representation and transferable prototypes
Junsong Fan, Yuxi Wang 0001, He Guan, Chunfeng Song, Zhaoxiang Zhang 0001 |
Frontiers Comput. Sci. | 3 |
| 2022 | MonoPoly: A practical monocular 3D object detector
He Guan, Chunfeng Song, Zhaoxiang Zhang 0001, Tieniu Tan |
Pattern Recognit. | 1 |
| 2018 | CMCGAN: A Uniform Framework for Cross-Modal Visual-Audio Mutual GenerationabstractVisual and audio modalities are two symbiotic modalities underlying videos, which contain both common and complementary information. If they can be mined and fused sufficiently, performances of related video tasks can be significantly enhanced. However, due to the environmental interference or sensor fault, sometimes, only one modality exists while the other is abandoned or missing. By recovering the missing modality from the existing one based on the common information shared between them and the prior information of the specific modality, great bonus will be gained for various vision tasks. In this paper, we propose a Cross-Modal Cycle Generative Adversarial Network (CMCGAN) to handle cross-modal visual-audio mutual generation. Specifically, CMCGAN is composed of four kinds of subnetworks: audio-to-visual, visual-to-audio, audio-to-audio and visual-to-visual subnetworks respectively, which are organized in a cycle architecture. CMCGAN has several remarkable advantages. Firstly, CMCGAN unifies visual-audio mutual generation into a common framework by a joint corresponding adversarial loss. Secondly, through introducing a latent vector with Gaussian distribution, CMCGAN can handle dimension and structure asymmetry over visual and audio modalities effectively. Thirdly, CMCGAN can be trained end-to-end to achieve better convenience. Benefiting from CMCGAN, we develop a dynamic multimodal classification network to handle the modality missing problem. Abundant experiments have been conducted and validate that CMCGAN obtains the state-of-the-art cross-modal visual-audio generation results. Furthermore, it is shown that the generated modality achieves comparable effects with those of original modality, which demonstrates the effectiveness and advantages of our proposed method. Zhaoxiang Zhang 0001, He Guan |
AAAI | 3 |
| 2018 | Integrating Both Visual and Audio Cues for Enhanced Video CaptionabstractVideo caption refers to generating a descriptive sentence for a specific short video clip automatically, which has achieved remarkable success recently. However, most of the existing methods focus more on visual information while ignoring the synchronized audio cues. We propose three multimodal deep fusion strategies to maximize the benefits of visual-audio resonance information. The first one explores the impact on cross-modalities feature fusion from low to high order. The second establishes the visual-audio short-term dependency by sharing weights of corresponding front-end networks. The third extends the temporal dependency to long-term through sharing multimodal memory across visual and audio modalities. Extensive experiments have validated the effectiveness of our three cross-modalities fusion strategies on two benchmark datasets, including Microsoft Research Video to Text (MSRVTT) and Microsoft Video Description (MSVD). It is worth mentioning that sharing weight can coordinate visual- audio feature fusion effectively and achieve the state-of-art performance on both BELU and METEOR metrics. Furthermore, we first propose a dynamic multimodal feature fusion framework to deal with the part modalities missing case. Experimental results demonstrate that even in the audio absence mode, we can still obtain comparable results with the aid of the additional audio modality inference module. Zhaoxiang Zhang 0001, He Guan |
AAAI | 3 |
| 2018 | Inception Donut Convolution for Top-down Semantic SegmentationabstractOne of recent trends in network architecture design confirms that the inception-block convolutional group is efficient, since it can aggregate spatial context information in lower dimensions without causing significant loss in representative capabilities. We believe that not only the strong correlation between adjacent cells, multi-scale feature extraction also plays a vital role in this novel module. In this paper, we extend the profits of the block to a top-down donut convolutional network for semantic segmentation task. Our network automatically learns rich convolution kernels to capture more structure prior. In the inception-block design, it overcomes the limitations in larger kernel size and adaptively captures different object-scales contexts without chain sampling. Our experiments demonstrate that the proposed inception-block donut convolutional network is orthogonal and can further improve the performance of most off-the-shelf bottom-up based methods. He Guan, Zhaoxiang Zhang 0001, Tieniu Tan |
ICPR | 1 |
| 2018 | Rethinking ReLU to Train Better CNNsabstractMost of convolutional neural networks share the same characteristic: each convolutional layer is followed by a nonlinear activation layer where Rectified Linear Unit (ReLU) is the most widely used. In this paper, we argue that the designed structure with the equal ratio between these two layers may not be the best choice since it could result in the poor generalization ability. Thus, we try to investigate a more suitable method on using ReL U to explore the better network architectures. Specifically, we propose a proportional module to keep the ratio between convolution and ReLU amount to be N:m (n>m). The proportional module can be applied in almost all networks with no extra computational cost to improve the performance. Comprehensive experimental results indicate that the proposed method achieves better performance on different benchmarks with different network architectures, thus verify the superiority of our work. Gangming Zhao, Zhaoxiang Zhang 0001, He Guan, Peng Tang 0005, Jingdong Wang 0001 |
ICPR | 3 |
| 2018 | View Decomposition and Adversarial for Semantic Segmentation
He Guan, Zhaoxiang Zhang 0001 |
PRICAI | 1 |