Yixin Chen 0004

dblp:59/983-4 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-2939-2541ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 8 since 2021Computer networks · 4 · 2 first-authorArtificial intelligence and machine learning · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Graph2Region: Efficient Graph Similarity Learning With Structure and Scale Restoration (Extended Abstract)
Zhouyang Liu, Yixin Chen 0004, Ning Liu 0015, Jiezhong He, Dongsheng Li 0001
ICDE2
2026 Hierarchy-Aware Neural Subgraph Matching with Enhanced Similarity Measure (Extended Abstract)
Zhouyang Liu, Ning Liu 0015, Yixin Chen 0004, Jiezhong He, Menghan Jia, Dongsheng Li 0001
ICDE3
2026 Rethinking Flexible Graph Similarity Computation: One-Step Alignment with Global Guidance
abstract
Graph Edit Distance (GED) is a widely used measure of graph similarity, valued for its flexibility in encoding domain knowledge through operation costs. However, existing learning-based approximation methods follow a modeling paradigm that decouples local candidate match selection from both operation costs and global dependencies between matches. This decoupling undermines their ability to capture the intrinsic flexibility of GED and often forces them to rely on costly iterative refinement to obtain accurate alignments. In this work, we revisit the formulation of GED and revise the prevailing paradigm, and propose Graph Edit Network (GEN), an implementation of the revised formulation that tightly integrates cost-aware expense estimation with globally guided one-step alignment. Specifically, GEN incorporates operation costs into node matching expenses estimation, ensuring match decisions respect the specified cost setting. Furthermore, GEN models match dependencies within and across graphs, capturing each match's impact on the overall alignment. These designs enable accurate GED approximation without iterative refinement. Extensive experiments on real-world and synthetic benchmarks demonstrate that GEN achieves up to a 37.8% reduction in GED predictive errors, while increasing inference throughput by up to 414x. These results highlight GEN's practical efficiency and the effectiveness of the revision. Beyond this implementation, our revision provides a principled framework for advancing learning-based GED approximation.
Zhouyang Liu, Ning Liu 0015, Yixin Chen 0004, Jiezhong He, Shuai Ma 0001, Dongsheng Li 0001
ICDE3
2025 TriFMatch: a flash subgraph matching algorithm with effective filtering techniques
Jiezhong He, Yixin Chen 0004, Menghan Jia, Zhouyang Liu, Dongsheng Li 0001, Kian-Lee Tan
Knowl. Inf. Syst.2
2025 Graph2Region: Efficient Graph Similarity Learning With Structure and Scale Restoration
Zhouyang Liu, Yixin Chen 0004, Ning Liu 0015, Jiezhong He, Dongsheng Li 0001
IEEE Trans. Knowl. Data Eng.2
2025 Hierarchy-Aware Neural Subgraph Matching With Enhanced Similarity Measure
abstract
Subgraph matching is challenging as it necessitates time-consuming combinatorial searches. Recent Graph Neural Network (GNN)-based approaches address this issue by employing GNN encoders to extract graph information and hinge distance measures to ensure containment constraints in the embedding space. These methods significantly shorten the response time, making them promising solutions for subgraph retrieval. However, they suffer from scale differences between graph pairs during encoding, as they focus on feature counts but overlook the relative positions of features within node-rooted subtrees, leading to disturbed containment constraints and false predictions. Additionally, their hinge distance measures lack discriminative power for matched graph pairs, hindering ranking applications. We propose NC-Iso, a novel GNN architecture for neural subgraph matching. NC-Iso preserves the relative positions of features by building the hierarchical dependencies between adjacent echelons within node-rooted subtrees, ensuring matched graph pairs maintain consistent hierarchies while complying with containment constraints in feature counts. To enhance the ranking ability for matched pairs, we introduce a novel similarity dominance ratio-enhanced measure, which quantifies the dominance of similarity over dissimilarity between graph pairs. Empirical results on nine datasets validate the effectiveness, generalization ability, scalability, and transferability of NC-Iso while maintaining time efficiency, offering a more discriminative neural subgraph matching solution for subgraph retrieval.
Zhouyang Liu, Ning Liu 0015, Yixin Chen 0004, Jiezhong He, Menghan Jia, Dongsheng Li 0001
IEEE Trans. Knowl. Data Eng.3
2024 Optimizing subgraph retrieval and matching with an efficient indexing scheme
Jiezhong He, Yixin Chen 0004, Zhouyang Liu, Dongsheng Li 0001
Knowl. Inf. Syst.2
2022 Fixed-Size Objects Encoding for Visual Relationship Detection
Hengyue Pan, Xin Niu 0002, Yixin Chen 0004, Peng Qiao, Zhen Huang 0006, Dongsheng Li 0001
Neural Process. Lett.4
2022 Online Learning Bipartite Matching with Non-stationary Distributions
abstract
Online bipartite matching has attracted wide interest since it can successfully model the popular online car-hailing problem and sharing economy. Existing works consider this problem under either adversary setting or i.i.d. setting. The former is too pessimistic to improve the performance in the general case; the latter is too optimistic to deal with the varying distribution of vertices. In this article, we initiate the study of the non-stationary online bipartite matching problem, which allows the distribution of vertices to vary with time and is more practical. We divide the non-stationary online bipartite matching problem into two subproblems, the matching problem and the selecting problem, and solve them individually. Combining Batch algorithms and deep Q-learning networks, we first construct a candidate algorithm set to solve the matching problem. For the selecting problem, we use a classical online learning algorithm, Exp3, as a selector algorithm and derive a theoretical bound. We further propose CDUCB as a selector algorithm by integrating distribution change detection into UCB. Rigorous theoretical analysis demonstrates that the performance of our proposed algorithms is no worse than that of any candidate algorithms in terms of competitive ratio. Finally, extensive experiments show that our proposed algorithms have much higher performance for the non-stationary online bipartite matching problem comparing to the state-of-the-art.
Jiaqi Zheng 0001, Guihai Chen, Yixin Chen 0004, Dongsheng Li 0001
ACM Trans. Knowl. Discov. Data5
2021 Employer-Employee Network for Conversational Recommendation
abstract
Traditional recommendation systems model user preferences based on past historical behaviors, thus unable to obtain dynamic user preferences. The conversational recommendation system (CRS) combines the conversational module with the recommendation module and overcomes the limitations by directly asking the user's preference for attributes. However, the existing CRS methods lack effective information propagation among various modules, making the model lack the basis for making correct decisions. In this paper, we propose an Employer-Employee network, which decomposes the actions into two stages, which are completed by two networks respectively. The Employer Network is responsible for analyzing information and making decisions (query or recommend), and the Employee Network is responsible for collecting information and performing tasks. Our contributions can be highlighted in three aspects: We first emphasize the importance of information propagation among multiple modules in the conversational recommendation system. Secondly, we propose an Employer-Employee (EE) network, which transforms each turn of action into a two-stage decision-making task handed over to two networks to complete. Thirdly, we conduct experiments on multiple datasets, and the experimental results show that our model achieves competitive performance compared with state-of-the-art baselines.
Zijing Yang, Xin Lin 0001, Liang He 0001, Yixin Chen 0004
IJCNN5
2021 Effective and Efficient Content Redundancy Detection of Web Videos
abstract
Currently, an unprecedentedly vast amount of videos are hosted on the Internet and shared by users across the world. Within these videos, a considerable portion is duplicate or near-duplicate. Consequently, building an effective yet efficient content-based redundancy detection system is of importance, as this research would be beneficial to a variety of applications. Despite the progress in this field, designing a practical detection system for web videos continues to be difficult, because of the contradictions between the accuracy and speed requirements. In this paper, we propose a novel near-duplicate video detection system, CompoundEyes, whose design philosophy deviates from the conventional feature-centered paradigm. Instead, the focus of our system has been shifted from the design of an advanced feature representation to the design of system architecture. This design methodology not only ensures a decent detection accuracy by the collaboration of the classifiers but also substantially accelerates the detection speed due to the low dimensionality of the feature representations and the exploitation of the parallelism among the components. Experiments have been conducted to demonstrate that the CompoundEyes is both accurate and fast.
Yixin Chen 0004, Dongsheng Li 0001, Yu Hua 0001, Wenbo He 0003
IEEE Trans. Big Data1
2020 ReLoca: Optimize Resource Allocation for Data-parallel Jobs using Deep Learning
abstract
Since under-allocating computation resource (e.g., CPU cores) causes suboptimal JCTs of data-parallel jobs, users are inclined to request excessive computation resource to decrease JCTs. However, over-allocating computation resource for data-parallel jobs incurs considerable system overheads (e.g., network communication and disk I/O overhead), which prolong the job completion time. In this paper, we propose ReLoca towards the optimal allocation of computation resource with the objective of minimizing the job completion time. ReLoca employs a deep neural network to guide the allocation of computation resource, by learning the impact of the operations in data-parallel jobs on the system overhead and computation time. Since training samples are time-consuming to collect, we develop an adaptive sampling method to preferably collect high-quality samples and thus overcome the issue of data scarcity. We apply ReLoca to improve Spark and conduct real experiments with five typical applications in big data analytics. Results show that ReLoca significantly reduces the average job completion time. Compared with the state-of-the-art method, ReLoca has higher prediction accuracy, needs fewer training samples and decreases the sampling overhead. With the prediction by ReLoca, the JCT decreases by 29.85%.
Zhiyao Hu, Dongsheng Li 0001, Dongxiang Zhang, Yixin Chen 0004
INFOCOM4
2020 Tag Pollution Detection in Web Videos via Cross-Modal Relevance Estimation
abstract
In the era of big data, web videos are known for their astronomical volume and great difficulties to be understood by computers. Therefore it is challenging to detect and curb tag pollution on social networking platforms. Intuitively, the pollution in the tags of a video can be identified by exploiting the tags of visually similar videos. From this intuition, we develop a semi-supervised approach to estimate the relevance between Internet videos and user-provided labels accurately and detect polluted labels accordingly. To further enhance the accuracy of relevance estimation and pollution detection, we introduce three multi-view multi-label models, which employ the coherence and differences between various similarity relations of videos. Compared with two quintessential multi-view fusion models, the proposed models consistently outperform or achieve comparable performance.
Yixin Chen 0004, Xinye Lin, Ke-shi Ge, Wenbo He 0003, Dongsheng Li 0001
IWQoS1
2016 CompoundEyes: Near-duplicate detection in large scale online video systems in the cloud
abstract
At the present time, billions of videos are hosted and shared in the cloud of which a sizable portion consists of near-duplicate video copies. An efficient and accurate content-based online near-duplicate video detection method is a fundamental research goal; as it would benefit applications such as duplication-aware storage, pirate video detection, polluted video tag detection, searching result diversification. Despite the recent progress made in near-duplicate video detection, it remains challenging to develop a practical detection system for large-scale applications that has good efficiency and accuracy performance. In this paper, we shift the focus from feature representation design to system design, and develop a novel system, called CompoundEyes, accordingly. The improvement in accuracy is achieved via well-organized classifiers instead of advanced feature design. Meanwhile, by applying simple features with reduced dimensionality and exploiting the parallelism of the detection architecture, we accelerate the detection speed. Through extensive experiments we demonstrate that the proposed detection system is accurate and fast. It takes approximately 1.45 seconds to process a video clip from a large video dataset, CC_WEB_VIDEO, with a 89% detection accuracy.
Yixin Chen 0004, Wenbo He 0003, Yu Hua 0001, Wen Wang 0018
INFOCOM1
2016 Cupid: Congestion-free consistent data plane update in software defined networks
abstract
With the popular applications of SDN in load balancing and failure recovery, the controller schedules affected flows to redundant paths to avoid network congestions and failures by updating flow tables in data plane. However, inconsistent flow table updating may lead to transient incorrect network behaviors or undesired performance degradation. Therefore, the consistency imposes dependencies among updates, so that the order of updates must be carefully considered to keep the consistency. To update flow tables consistently and efficiently, in this paper, we propose an update ordering approach — Cupid. To avoid high overhead in update ordering, we divide the global dependencies among updates into local restrictions by: 1) partitioning a new routing path into several independent segments, 2) identifying critical nodes controlling traffic shifting between the old path and new path, and 3) constructing a dependency graph among critical nodes for potential congested links. We then design a heuristic algorithm to resolve the dependency graph. To save the flow table space, a switch keeps only one flow entry with multiple ports for a flow during updating. Our simulation shows that Cupid schedules updates at least 2 times faster and has less throughput losses than the state-of-the-art approaches in both fat-tree and mesh networks.
Wen Wang 0018, Wenbo He 0003, Jinshu Su, Yixin Chen 0004
INFOCOM4
2015 Marlin: Taming the big streaming data in large scale video similarity search
abstract
The extreme volume and staggeringly increasing rate inevitably produce unprecedented pressure on any large scale video sharing and hosting systems. Among the efforts to mitigate this pressure, content-based video similarity search is becoming more and more important with the exponential growth of the data size. Though various approaches have been proposed to address this problem, they are mainly focusing on the retrieval accuracy thus bringing video features with high complexity. Due to the complexity of the feature, these systems are based on the assumption that features representing videos have been obtained offline and stored in the database statically. However, the on-call efforts to move the feature extraction and similarity search from offline to online have been ignored in previous work. In this paper, we propose Marlin, a streaming data processing pipeline that efficiently extracts video features and retrieves video similarity information in a large scale video data system. We design a streaming feature extractor to handle the videos streaming into the system and establish the fined-grained resource allocation with a resource-aware data abstraction layer over streaming data to allocate computing resources among the videos with various resource demands. Besides that, we are pipelining the feature extraction and similarity search process with a distributed feature index, which supports real-time query and incremental index update. The experimental and the extensive real-world workload driven simulation results show that the proposed stream processing architecture achieves 25X speedup against the sequential feature extraction algorithm and 23X speedup against the sequential similarity search with a subsecond similarity query latency for a single request.
Wenbo He 0003, Yu Hua 0001, Yixin Chen 0004
IEEE BigData4