Yibin Chen

dblp:03/139 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 1 since 2021Systems, architecture and hardware · 2 · 2 first-authorComputer networks · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SAT: Balancing Reasoning Accuracy and Efficiency with Stepwise Adaptive Thinking
abstract
Weiyang Huang, Xuefeng Bai, Kehai Chen, Xinyang Chen, Yibin Chen, Weili Guan, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Weiyang Huang, Xuefeng Bai 0001, Kehai Chen, Xinyang Chen 0001, Yibin Chen, Weili Guan, Min Zhang 0005
ACL (1)5
2026 MFDSGC: a novel pedestrian crossing intention prediction method for autonomous vehicles
Jingwei Cao, Guoyang Hou, Jiawang Lv, Yibin Chen, Liming Di
Adv. Eng. Informatics4
2025 SheetAgent: Towards a Generalist Agent for Spreadsheet Reasoning and Manipulation via Large Language Models
abstract
Spreadsheets are ubiquitous across the World Wide Web, playing a critical role in enhancing work efficiency across various domains. Large language model (LLM) has been recently attempted for automatic spreadsheet manipulation but has not yet been investigated in complicated and realistic tasks where reasoning challenges exist (e.g., long horizon manipulation with multi-step reasoning and ambiguous requirements). To bridge the gap with the real-world requirements, we introduce SheetRM, a benchmark featuring long-horizon and multi-category tasks with reasoning-dependent manipulation caused by real-life challenges. To mitigate the above challenges, we further propose SheetAgent, a novel autonomous agent that utilizes the power of LLMs. SheetAgent consists of three collaborative modules: Planner, Informer, and Retriever, achieving both advanced reasoning and accurate manipulation over spreadsheets without human interaction through iterative task reasoning and reflection. Extensive experiments demonstrate that SheetAgent delivers 20--40% pass rate improvements on multiple benchmarks over baselines, achieving enhanced precision in spreadsheet manipulation and demonstrating superior table reasoning abilities. More details and visualizations are available at the https://sheetagent.github.io/. The datasets and source code are available at https://anonymous.4open.science/r/SheetAgent.
Yibin Chen, Yifu Yuan, Yan Zheng 0002, Jinyi Liu 0002, Fei Ni 0001, Jianye Hao, Hangyu Mao
WWW1
2024 Take Off the Training Wheels! Progressive In-Context Learning for Effective Alignment
abstract
Recent studies have explored the working mechanisms of In-Context Learning (ICL).However, they mainly focus on classification and simple generation tasks, limiting their broader application to more complex generation tasks in practice.To address this gap, we investigate the impact of demonstrations on token representations within the practical alignment tasks.We find that the transformer embeds the task function learned from demonstrations into the separator token representation, which plays an important role in the generation of prior response tokens.Once the prior response tokens are determined, the demonstrations become redundant.Motivated by this finding, we propose an efficient Progressive In-Context Alignment (PICA) method consisting of two stages.In the first few-shot stage, the model generates several prior response tokens via standard ICL while concurrently extracting the ICL vector that stores the task function from the separator token representation.In the following zero-shot stage, this ICL vector guides the model to generate responses without further demonstrations.Extensive experiments demonstrate that our PICA not only surpasses vanilla ICL but also achieves comparable performance to other alignment tuning methods.The proposed training-free method reduces the time cost (e.g., 5.45×) with improved alignment performance (e.g., 6.57+).Consequently, our work highlights the application of ICL for alignment and calls for a deeper understanding of ICL for complex generations.
Dongfang Li 0002, Xinshuo Hu, Xinping Zhao, Yibin Chen, Baotian Hu, Min Zhang 0005
EMNLP5
2024 SEER: Self-Aligned Evidence Extraction for Retrieval-Augmented Generation
abstract
Recent studies in Retrieval-Augmented Generation (RAG) have investigated extracting evidence from retrieved passages to reduce computational costs and enhance the final RAG performance, yet it remains challenging.Existing methods heavily rely on heuristic-based augmentation, encountering several issues: (1) Poor generalization due to hand-crafted context filtering; (2) Semantics deficiency due to rulebased context chunking; (3) Skewed length due to sentence-wise filter learning.To address these issues, we propose a model-based evidence extraction learning framework, SEER, optimizing a vanilla model as an evidence extractor with desired properties through selfaligned learning.Extensive experiments show that our method largely improves the final RAG performance, enhances the faithfulness, helpfulness, and conciseness of the extracted evidence, and reduces the evidence length by 9.25 times.The code will be available at https://github.com/HITsz-TMG/SEER.
Xinping Zhao, Dongfang Li 0002, Boren Hu, Yibin Chen, Baotian Hu, Min Zhang 0005
EMNLP5
2024 A small object detection network for remote sensing based on CS-PANet and DSAN
Lei Zhang 0001, Fengxian Wang, Yibin Chen
Multim. Tools Appl.6
2023 WAG-NAT: Window Attention and Generator Based Non-Autoregressive Transformer for Time Series Forecasting
Yibin Chen, Ailan Xu, Qiang Sun 0001, Chen Xu 0005
ICANN (6)1
2023 Evaluation of Reachability Queries Based on Recursive DAG Decomposition
abstract
Let$G(V,\,E)$be a digraph (directed graph) with$n$nodes and$e$arcs. Digraph$G^{\ast}=(V,E^{\ast})$) is the reflexive, transitive closure if$v\rightarrow u\, \epsilon\, E^{\ast}$iff there is a path from$v$to$u$in$G$. Efficient storage of$G^{\ast}$is important for supporting reachability queries which are not only common in graph databases, but also serve as fundamental operations used in many graph algorithms. A lot of strategies have been proposed based on the graph labeling, by which each node is assigned with certain labels such that the reachability of any two nodes through a path can be determined by their labels. Among them are interval labeling, chain decomposition, 2-hop labeling, and path-trees, as well as partial index based methods. However, due to the very large size of many real world graphs, the computational cost and size of labels using existing methods would prove too expensive to be practical. In this paper, we propose a new approach to deduct and decompose a graph into a series of spanning trees and transform a query$q$to a series of subqueries each evaluated against a spanning tree. Using the so-called tree labeling, each subquery needs only O(1) time. More importantly, the number of such subqueries is$\ll_{n}$. Thus,$q$can be evaluated very efficiently. We demonstrate both analytically and empirically the efficiency and effectiveness of our method. While the query time of our method is orders of magnitude better than almost all the existing strategies, its indexing time and index sizes are comparable to them.
Yangjun Chen, Yibin Chen
IEEE Trans. Knowl. Data Eng.2
2019 An industrial missing values processing method based on generating model
Zhaolin Yuan, Yibin Chen, Bingyang Shen, Aixiang Wu
Comput. Networks3
2019 Introducing Cuts Into a Top-Down Process for Checking Tree Inclusion
abstract
By the ordered tree inclusion we will check whether a pattern tree P can be included in a target tree T, where the order of siblings in both Pand T matters. This problem has many applications in practice, such as retrieval of documents, data mining, and RNA structure matching. In this paper, we propose an efficient algorithm for this problem. Its time complexity is bounded by O(|T| · min{hP; |leaves(P)|}), with O(|T| + |P|) space being used, where hP(hT) represents the height of P (resp., T) and leaves (P) stands for the set of the leaves of P. Up to now the best algorithm for this problem needs Θ(|T| · |leaves(P)|) time and O(|P| + |T|) space. Extensive experiments have been done, which show that the new algorithm can perform much better than the existing ones in practice.
Yangjun Chen, Yibin Chen
IEEE Trans. Knowl. Data Eng.2
2018 Cross Modal Multiscale Fusion Net for Real-time RGB-D Detection
abstract
This paper presents a novel multi-modal CNN architecture for object detection by exploiting complementary input cues in addition to sole color information. Our one-stage architecture fuses the multiscale mid-level features from two individual feature extractor, so that our end-to-end net can accept cross modal streams to obtain high-precision detection results. In comparison to other cross modal fusion neural networks, our solution successfully reduces runtime to meet the real-time requirement with still high-level accuracy. Experimental evaluation on challenging NYUD2 dataset shows that our network achieves 49.1% mAP, and processes images in real-time at 35.3 frames per second on one single Nvidia GTX1080 GPU. Compared to baseline one stage network SSD on RGB images which gets 39.2% mAP, our method has great accuracy improvement.
Kejie Yin, Sheng Liu 0002, Ruyu Liu, Yibin Chen
ICPR4
2018 A Novel Visual Tracking Method Based on Moth-Flame Optimization Algorithm
Huanlong Zhang, Xiujiao Zhang, Xiaoliang Qian, Yibin Chen
PRCV (4)4
2016 Dynamic Power Allocation for a Hybrid Energy Harvesting Transmitter with Multiuser in Fading Channels
abstract
In this work, we consider a multiuser communication system in fading channels where the transmitter is supplied by hybrid energy sources including power grid and various renewable sources. Specially, the energy harvested from various renewable sources is stored in a limited capacity buffer, and the joint energy incoming is time-varying and possibly unpredictable. In addition, data arrives randomly to the transmitter and queues according to the individual receivers, the wireless channels fluctuate randomly due to fading. Our goal is, under this condition to develop a dynamic power allocation algorithm so as to minimize the time average amount of energy consumed from the power grid over an infinite horizon, subjecting to all data queues stability. The issue is formulated as a stochastic optimization problem and solved by Lyapunov optimization technique which does not require the statistical probabilities of energy harvesting process, data arrivals process and channel state. Simulation results demonstrate that our proposed algorithm provides obviously better performance than other two simple greedy algorithms, meanwhile the algorithm gives a guarantee that the maximum delay of all data queues cannot exceed a given value.
Didi Liu, Jiming Lin, Yibin Chen
VTC Fall5
2015 Tree Inclusion Checking Revisited
Yangjun Chen, Yibin Chen
DATA2
2013 Layered moving-object segmentation for stereoscopic video using motion and depth information
Yibin Chen, Canhui Cai, Kai-Kuang Ma, Xiaolan Wang 0006
J. Vis. Commun. Image Represent.1
2011 Decomposing DAGs into spanning trees: A new way to compress transitive closures
abstract
Let G(V, E) be a digraph (directed graph) with n nodes and e edges. Digraph G* = (V, E*) is the reflexive, transitive closure if (v, u) ∈ E* iff there is a path from v to u in G. Efficient storage of G* is important for supporting reachability queries which are not only common on graph databases, but also serve as fundamental operations used in many graph algorithms. A lot of strategies have been suggested based on the graph labeling, by which each node is assigned with certain labels such that the reachability of any two nodes through a path can be determined by their labels. Among them are interval labelling, chain decomposition, and 2-hop labeling. However, due to the very large size of many real world graphs, the computational cost and size of labels using existing methods would prove too expensive to be practical. In this paper, we propose a new approach to decompose a graph into a series of spanning trees which may share common edges, to transform a reachability query over a graph into a set of queries over trees. We demonstrate both analytically and empirically the efficiency and effectiveness of our method.
Yangjun Chen, Yibin Chen
ICDE2
2010 Histogram-offset-based color correction for multi-view video coding
abstract
In multi-view video system, variations of different camera setups (for example, camera positions, lighting conditions and camera characteristics) might cause discrepancies on the luminance and chrominance components of different views. From the viewpoint of source compression, this will lead to inaccurate inter-view prediction and lower coding efficiency. In this paper, a histogram-offset-based color correction method is developed for benefitting multi-view video coding. First, disparity estimation is conducted on the rank-transformed domain to identify the maximum matching regions between the reference view and the target view. Within the identified matching regions, the histograms of the reference view and the target view are then calculated, respectively. By using an iterative thresholding approach, a histogram offset is generated and exploited to correct the target view. Experimental results have shown that the proposed color correction method outperforms the histogram matching method on the improvement of coding efficiency.
Yibin Chen, Kai-Kuang Ma, Canhui Cai
ICIP1
2010 Automated Design Debugging With Maximum Satisfiability
abstract
As contemporary very large scale integration designs grow in complexity, design debugging has rapidly established itself as one of the largest bottlenecks in the design cycle today. Automated debug solutions such as those based on Boolean satisfiability (SAT) enable engineers to reduce the debug effort by localizing possible error sources in the design. Unfortunately, adaptation of these techniques to industrial designs is still limited by the performance and capacity of the underlying engines. This paper presents a novel formulation of the debugging problem using MaxSAT to improve the performance and applicability of automated debuggers. Our technique not only identifies errors in the design but also indicates when the bug is excited in the error trace. MaxSAT allows for a simpler formulation of the debugging problem, reducing the problem size by 80% compared to a conventional SAT-based technique. Empirical results demonstrate the effectiveness of the proposed formulation as run-time improvements of 4.5 × are observed on average. This paper introduces two performance improvements to further reduce the time required to find all error sources within the design by an order of magnitude.
Yibin Chen, Sean Safarpour, João Marques-Silva 0001, Andreas G. Veneris
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2009 Spatial and temporal design debug using partial MaxSAT
abstract
Design debug remains one of the major bottlenecks in the VLSI design cycle today. Existing automated solutions strive to aid engineers in reducing the debug effort by identifying possible error sources in the design. Unfortunately, these techniques do not provide any information regarding the time at which the bug is active during an error trace or counter-example. This work introduces an automated debug technique that provides the user with both spatial and temporal information about the source of error. The proposed method is based on a Partial MaxSAT formulation which models errors at the CNF clause level instead of the traditional gate or module level. Thus, error sites are identified based on erroneous implications that correspond to locations both in the design and in the error trace. Experiments demonstrate that we can provide this additional information at no extra cost in run time and are able to prune about 61% of all simulation time frames from the debugging process. When compared to a trivial formulation we observe a performance improvement of up to two orders of magnitude and 5× on average when using the proposed formulation.
Yibin Chen, Sean Safarpour, Andreas G. Veneris, João Marques-Silva 0001
ACM Great Lakes Symposium on VLSI1
2009 Stereoscopic video error concealment for missing frame recovery using disparity-based frame difference projection
abstract
At low bit-rate video communications, packet loss may easily cause whole-frame loss that, in return, leads to annoying frame drop phenomenon. In this paper, a novel error concealment algorithm is specifically developed for stereoscopic video, called the disparity-based frame difference projection (DFDP), to recover the lost frames at the decoder. The proposed DFDP contains three key components: 1) change detection, 2) disparity estimation, and 3) frame difference projection, which exploits both the intra-view frame difference from one view and interview correlation to estimate the lost frame in another view. The change region computed on the correctly received frame will be used to predict the change region between current missing frame and its previous frame through the estimated disparity, which is the summation of the estimated global disparity and the estimated local disparity. Experimental results have shown that the proposed stereoscopic video error concealment method can effectively restore the lost frames at the decoder and deliver attractive performance, in terms of objective measurement (in peak signal-to-noise ratio) and subjective visual quality.
Yibin Chen, Canhui Cai, Kai-Kuang Ma
ICIP1
2008 An Efficient Algorithm for Answering Graph Reachability Queries
abstract
Given a directed graph G, to check whether a node v is reachable from another node u through a path is often required. In a database system, such an operation is called a recursion computation or reachability checking and not efficiently supported. The reason for this is that the space to store the whole transitive closure of G is prohibitively high. In this paper, we address this issue and propose an 0(n2+ bnradic(b)) time algorithm to decompose a directed acyclic graph (DAG) into a minimized set of disjoint chains to facilitate reachability checking, where n is the number of the nodes and b is the DAG's width, defined to be the size of a largest node subset U of the DAG such that for every pair of nodes u, v isin U, there does not exist a path from u to v or from v to u. Using this algorithm, we are able to label a graph in 0(be) time and store all the labels in O(bn) space with O(logb) reachability checking time, where e is the number of the edges of the DAG. The method can also be extended to handle cyclic directed graphs. Experiments have been performed, showing that our method is promising.
Yangjun Chen, Yibin Chen
ICDE2
2006 A new tree inclusion algorithm
Yangjun Chen, Yibin Chen
Inf. Process. Lett.2
2006 On the Signature Tree Construction and Analysis
abstract
Advanced database application areas, such as computer aided design, office automation, digital libraries, data-mining, as well as hypertext and multimedia systems, need to handle complex data structures with set-valued attributes, which can be represented as bit strings, called signatures. A set of signatures can be stored in a file, called a signature file. In this paper, we propose a new method to organize a signature file into a tree structure, called a signature tree, to speed up the signature file scanning and query evaluation. In addition, the average time complexity of searching a signature tree is analyzed and how to maintain a signature tree on disk is discussed. We also conducted experiments, which show that the approach of signature trees provides a promising index structure
Yangjun Chen, Yibin Chen
IEEE Trans. Knowl. Data Eng.2