Huyin Zhang

dblp:53/175 · DBLP profile ↗
← Back
38ranked-venue papers
2as first author
25since 2021 · last 2026
0000-0003-3820-3586ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 12 since 2021Systems, architecture and hardware · 11 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 since 2021Computer networks · 4 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 WorldSnake: Semantic-geometric decoupling for open-vocabulary contour-based instance segmentation
Yejun Wu, Jiao Zhan, Chi Guo, Huyin Zhang
Knowl. Based Syst.4
2025 SAMSnake: A generic contour-based instance segmentation network assisted by Efficient Segment Anything Model
Yejun Wu, Jiao Zhan, Chi Guo, Huyin Zhang
Neural Networks4
2024 Contextual Augmented Global Contrast for Multimodal Intent Recognition
abstract
Multimodal intent recognition (MIR) aims to perceive the human intent polarity via language, visual, and acoustic modalities. The inherent intent ambiguity makes it challenging to recognize in multimodal scenarios. Existing MIR methods tend to model the individual video independently, ignoring global contextual information across videos. This learning manner inevitably introduces perception biases, exacerbated by the inconsistencies of the multimodal representation, amplifying the intent uncertainty. This challenge motivates us to explore effective global context modeling. Thus, we propose a context-augmented global contrast (CAGC) method to capture rich global context features by mining both intra-and cross-video context interactions for MIR. Concretely, we design a context-augmented transformer module to extract global context dependencies across videos. To further alleviate error accumulation and interference, we develop a cross-video bank that retrieves effective video sources by considering both intentional tendency and video similarity. Furthermore, we introduce a global context-guided contrastive learning scheme, designed to mitigate inconsistencies arising from global context and individual modalities in different feature spaces. This scheme incorporates global cues as the supervision to capture robust the multimodal intent representation. Experiments demonstrate CAGC obtains superior performance than state-of-the-art MIR methods. We also generalize our approach to a closely related task, multimodal sentiment analysis, achieving the comparable performance.
Kaili Sun, Zhiwen Xie, Mang Ye, Huyin Zhang
CVPR4
2024 SDGIN: Structure-aware dual-level graph interactive network with semantic roles for visual dialog
Kaili Sun, Zhiwen Xie, Chi Guo, Huyin Zhang
Knowl. Based Syst.4
2024 Fast and isolation guaranteed coflow scheduling via traffic forecasting in multi-tenant environment
Chenghao Li 0001, Huyin Zhang, Sheng Hao 0001
J. Supercomput.2
2024 Simple hierarchical PageRank graph neural networks
Huyin Zhang, Shiming Tao, Xiying Fan
J. Supercomput.2
2024 Learning Heterogeneous Relation Graph and Value Regularization Policy for Visual Navigation
abstract
The goal of visual navigation is steering an agent to find a given target object with current observation. It is crucial to learn an informative visual representation and robust navigation policy in this task. Aiming to promote these two parts, we propose three complementary techniques, heterogeneous relation graph (HRG), a value regularized navigation policy (VRP), and gradient-based meta learning (ML). HRG integrates object relationships, including object semantic closeness and spatial directions, e.g., a knife is usually co-occurrence with bowl semantically or located at the left of the fork spatially. It improves visual representation learning. Both VRP and gradient-based ML improve robust navigation policy, regulating this process of the agent to escape from the deadlock states such as being stuck or looping. Specifically, gradient-based ML is a type of supervision method used in policy network training, which eliminates the gap between the seen and unseen environment distributions. In this process, VRP maximizes the transformation of the mutual information between visual observation and navigation policy, thus improving more informed navigation decisions. Our framework shows superior performance over the current state-of-the-art (SOTA) in terms of success rate and success weighted by length (SPL). Our HRG outperforms the Visual Genome knowledge graph on cross-scene generalization with$\approx$$56\%$and$\approx$$39\%$improvement on Hits@$5^{*}$(proportion of correct entities ranked in top 5) and MRR$^{*}$(mean reciprocal rank), respectively. Our code and HRG datasets will be made publicly available in the scientific community.
Kang Zhou 0003, Chi Guo, Wenfei Guo, Huyin Zhang
IEEE Trans. Neural Networks Learn. Syst.4
2023 Optimal Graph Transformer Viterbi knowledge inference network for more successful visual navigation
Kang Zhou 0003, Chi Guo, Huyin Zhang, Bohan Yang 0008
Adv. Eng. Informatics3
2023 Urban ride-hailing demand prediction with multi-view information fusion deep learning framework
Yonghao Wu, Huyin Zhang, Shiming Tao
Appl. Intell.2
2023 Theoretical modeling and analysis of Uplink performance for Three-dimension spatial RISs-aided Wireless communication systems
Sheng Hao 0001, Huyin Zhang, Jianqun Cui
Comput. Commun.3
2023 Large-capacity image data hiding based on table look-up
Wenjia Ding, Huyin Zhang, Ralf Reulke
Multim. Tools Appl.2
2022 Simplified multilayer graph convolutional networks with dropout
Huyin Zhang, Shiming Tao
Appl. Intell.2
2022 Graph representation learning via simple jumping knowledge networks
Huyin Zhang, Shiming Tao, Sheng Hao 0001
Appl. Intell.2
2022 Improving indoor visual navigation generalization with scene priors and Markov relational reasoning
Kang Zhou 0003, Chi Guo, Huyin Zhang
Appl. Intell.3
2022 Semi-supervised classification via full-graph attention neural networks
Huyin Zhang, Shiming Tao
Neurocomputing2
2022 HVLM: Exploring Human-Like Visual Cognition and Language-Memory Network for Visual Dialog
Kaili Sun, Chi Guo, Huyin Zhang
Inf. Process. Manag.3
2022 Reversible image data hiding based on scalable difference expansion
Wenjia Ding, Huyin Zhang, Ralf Reulke
Pattern Recognit. Lett.2
2022 Syntax-Aware graph convolutional network for the recognition of chinese implicit inter-sentence relations
Kaili Sun, Huyin Zhang, Chi Guo, Linfei Yuan, Quan Hu
J. Supercomput.3
2022 MVDLSTM: MultiView deep LSTM framework for online ride-hailing order prediction
Yonghao Wu, Huyin Zhang, Shiming Tao
J. Supercomput.2
2022 Travel order quantity prediction via attention-based bidirectional LSTM networks
Huyin Zhang, Shiming Tao
J. Supercomput.2
2022 Relational attention-based Markov logic network for visual navigation
Kang Zhou 0003, Chi Guo, Huyin Zhang
J. Supercomput.3
2021 Fusing Semantic Segmentation and Object Detection for Visual SLAM in Dynamic Scenes
abstract
The assumption of static scenes limits the performance of traditional visual SLAM. Many existing solutions adopt deep learning methods or geometric constraints to solve the problem of dynamic scenes, but these schemes are either low efficiency or lack of robustness to a certain extent. In this paper, we propose a solution combining object detection and semantic segmentation to obtain the prior contours of potential dynamic objects. With this prior information, geometric constraints techniques are utilized to assist with removing dynamic feature points. Finally, the evaluation with the public datasets demonstrates that our proposed method can improve the accuracy of pose estimation and robustness of visual SLAM with no efficiency loss in high dynamic scenarios.
Chi Guo, Huyin Zhang
VRST4
2021 Multi goals and multi scenes visual mapless navigation in indoor using meta-learning and scene priors
Chi Guo, Binhan Luo, Huyin Zhang
Neurocomputing4
2021 Fair and near-optimal coflow scheduling without prior knowledge of coflow size
Chenghao Li 0001, Huyin Zhang, Wenjia Ding
J. Supercomput.2
2021 Intelligibility Enhancement Via Normal-to-Lombard Speech Conversion With Long Short-Term Memory Network and Bayesian Gaussian Mixture Model
abstract
Speech communications and interactions frequently occur in a variety of environments. Noise in the environment significantly degrades speech intelligibility when speaking and listening. Especially in the listening stage, even if the multimedia terminal outputs clean speech, it is still difficult for listeners to obtain information. Intelligibility enhancement (IENH) of speech is a technique for overcoming the environmental noise in the listening stage. It implements a perceptual enhancement of non-noisy speech. This study focuses on IENH via normal-to-Lombard speech conversion, inspired by a well known acoustic mechanism named the Lombard effect. Our method combines the long short-term memory (LSTM) network and Bayesian Gaussian mixture model (BGMM) to build a conversion architecture. Compared with baselines, it has three main advantages: 1) an LSTM network is used for spectral tilt mapping with fully considering short-term correlations and high-dimensional expression abilities; 2) the aperiodicity (AP) is mapped together with the fundamental frequency ($F_0$) by a BGMM, which considers their relevance constraints and the importance of APs; 3) the gender-dependent mapping is used for$F_0$and APs to consider distribution differences between genders. Experiments indicate that our method gets better performance in both objective and subjective tests.
Xiaochen Wang 0001, Ruimin Hu, Huyin Zhang, Shanfa Ke
IEEE Trans. Multim.4
2020 Normal-To-Lombard Speech Conversion by LSTM Network and BGMM for Intelligibility Enhancement of Telephone Speech
abstract
Noise in the environment significantly decreases the speech intelligibility of telephone conversations. Despite clean speech output from the device, the listener is still hard to get information. This study focuses on intelligibility enhancement (IENH) of telephone speech in near-end background noise based on normal-to-Lombard speech conversion. The proposed approach uses long short-term memory (LSTM) and Bayesian Gaussian mixture model (BGMM) to build the speech mapping model. Compared with previous studies, we fully consider the short-term correlations of speech and implement feature mappings with higher dimensional features and more types of features. Evaluations indicate that the proposed approach has achieved better results in both objective and subjective evaluation.
Xiaochen Wang 0001, Ruimin Hu, Huyin Zhang, Shanfa Ke
ICME4
2020 Delay-tolerant routing and message scheduling for CR-VANETs
Jing Wang 0063, Huyin Zhang, Xing Tang 0001, Zongpeng Li
Future Gener. Comput. Syst.2
2020 A comprehensive approach for road marking detection and recognition
Huyin Zhang, Jinsheng Xiao, Shejie Lu, Reinhard Klette, Mohammad Norouzifard
Multim. Tools Appl.2
2020 Theoretical modeling for performance analysis of IEEE 1901 power-line communication networks in the multi-hop environment
Sheng Hao 0001, Huyin Zhang
J. Supercomput.2
2019 Coflow scheduling algorithm based density peaks clustering
Chenghao Li 0001, Huyin Zhang
Future Gener. Comput. Syst.2
2019 An improved image mixed noise removal algorithm based on super-resolution algorithm and CNN
Huyin Zhang, Jinsheng Xiao, Shejie Lu, Mohammad Norouzifard
Neural Comput. Appl.2
2018 Graph cut based clustering for cognitive radio ad hoc networks without common control channels
Huyin Zhang, Fang Xu 0001
Wirel. Networks1
2017 Virtualized Resource Sharing in Cloud Radio Access Networks Through Truthful Mechanisms
abstract
In the recent paradigm of cloud radio access networks (C-RAN), signal processing functions at the base stations (BSs) are virtualized and migrated into a mobile cloud that maintains a pool of virtual BS (VBS) instances. Remote radio heads and antennae at the BSs are connected to the VBS pool by fronthaul fiber links. Mobile operators may lease resources from the tower company who owns the C-RAN infrastructure. We study auction mechanisms for efficiently sharing C-RAN resources among mobile operators. Leveraging randomized rounding, we design an offline C-RAN auction mechanism that can achieve truthfulness and near-optimal social welfare. For the more realistic setting of online bid arrival, we design an online algorithm that executes in polynomial time and achieves a competitive ratio of $(1-\epsilon )$ . A tailored fractional Vickrey-Clarke-Groves mechanism works in concert with the online algorithm to elicit truthful bids. Extensive simulation studies verify the efficacy of our C-RAN auction mechanisms.
Sijia Gu, Zongpeng Li, Chuan Wu 0001, Huyin Zhang
IEEE Trans. Commun.4
2016 Delay Tolerant Routing for Cognitive Radio Vehicular Ad Hoc Networks
abstract
Cognitive Radio Vehicular Ad Hoc Networks (CRVANETs) can resolve the conflicts between limited spectrum resource and vehicular communication service needs, and thus it receives much concern. Most existing routing solutions for Cognitive Radio Ad Hoc Networks (CRAHNs) or VANETs cannot be directly applied to CR-VANETs due to the high mobility of vehicles and varying availability of spectrums. To fill this gap, with the proliferation of delay tolerant applications, this paper presents a delay tolerant routing strategy for CR-VANETs with the object of maximizing the packet delivery ratio from the source to the destination. The proposed routing scheme is composed of two parts: the routing policy and the message replication policy. In the first policy, we select a concurrent forwarding set under the restriction of CR spectrum holes, and choose an optimized routing path for the next hop. In the second policy, we propose a hybrid and adaptable method relying on Drop-Oldest and Drop-Most strategies, to reduce the delivery cost. Through extensive simulations, we demonstrate that the proposed delay tolerant routing mechanism provides with better packet delivery ratio compared to state-of-the-art protocols.
Jing Wang 0063, Huyin Zhang, Xing Tang 0001
ICPADS2
2016 Spectrum-Aware Clustering with Proactive Handoff for Distributed Cognitive Radio Ad Hoc Networks
Huyin Zhang, Fang Xu 0001
WASA1
2008 MEANS: A Micro-thrEad Architecture for Network Server
abstract
Internet applications require high-performance network server architecture. This paper proposes a kind of software architecture for network server, MEANS, which aims at supporting Internet applications. By introducing a new thread abstraction, micro-thread, MEANS upwardly provides the micro-thread environment to programmers, and downwardly accesses the OS services concurrently through traditional thread primitives, and uses the event-driven mechanism to manage and schedule micro-threads, which takes advantage of both multithread and event-driven architecture. Moreover, MEANS is general, scalable, robust and adaptable. By preliminary evaluation, in terms of the concurrence policy, MEANS has features similar to event-driven architecture; and it outperforms multithread architecture in I/O accessing.
Yingchun Lei, Yili Gong, Huyin Zhang
PDP4
2006 An Overtime-Tolerance Strategy for Advance Reservation
abstract
In advance reservation environment, failure discovery strategy is very important to keep the normal running of whole system. One job's overtime (exceeding its booked time) may lead to a serious of jobs' abnormal termination in co-allocation environment. So there should be some strategies to solve this problem. In this paper, a novel overtime-tolerance strategy is introduced. The strategy can tolerate active job exceeding its booking time and redirect transparently later inactive jobs to other resources. The architecture of implementing our strategy is described in this paper too. Simulation shows our overtime-tolerance strategy can decrease the abnormal termination ratio and improve the system throughput
Chanle Wu, Jianqun Cui, Huyin Zhang, Gang Ye
PDCAT4
2005 Heterogeneous Resource Selection with Portfolio Optimization in Grid Economy
abstract
The problem of resource management and scheduling has always been the main challenge in grid. Grid economy is looked on as a good solution for obtaining better performance than other methods for resource management in dynamic grid environments. But, the decision making about resources selection with portfolio optimization has been paid little attention to in grid economy. This paper introduced the concept and methods with portfolio optimization in grid economy. Keywords: Grid Economy, Portfolio Optimization, Linear Programming, Max-Min Portfolio
Yu Hua 0001, Chanle Wu, Huyin Zhang
PDCAT3