VLDB 2026 Research / reviewers in the wild / expert
Yuzhe Jin
dblp:05/7825
· DBLP profile ↗
16ranked-venue papers
10as first author
2since 2021 · last 2026
0009-0009-4875-0226ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 4Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 3 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Theory of computation · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UntCC: Untangling Composite Commits Using Structural and Semantic InformationabstractSmall and focused commits are highly valued in modern software development. However, developers sometimes submit a commit with more than one concern, represented by several lines of code changes for a specific purpose, e.g., adding new features or fixing bugs. Such composite commits confuse developers during code reviews as well as other software activities, resulting in various issues. Existing studies predominantly leverage code structure to untangle composite commits, but without considering code semantics that have been demonstrated to be important in many related studies. In this article, we propose UNTCC, a new approach that uses structural and semantic information forUNTanglingCompositeCommits. To achieve structural information, we propose the code change graph, a fine-grained, text-attributed graph representation of a commit, incorporating before-change and after-change code dependencies; and UNTCC employs the graph autoencoder to learn its structural representation. To achieve semantic information, UNTCC leverages a large language model (Llama-3.2-3B) to learn joint embeddings of the raw commit and its aligned graph representation, which guide the division of different concerns within a composite commit. The experimental evaluation using 27,853 composite commits from 9 C# and 10 Java projects shows that in terms of Accuracya/Accuracyc, UNTCC achieves 94%/74% in C# and 77%/54% in Java, outperforming state-of-the-art approaches by 2%—623%/32%—573% in C# and 22%—285%/35%—286% in Java. The results indicate that UNTCC can effectively untangle composite commits. Yuzhe Jin, Lanxin Yang, He Zhang 0001, Gongyuan Li, Bohan Liu 0003, Xin Zhou 0016, Hongyu Kuang, Liming Dong 0001 |
IEEE Trans. Software Eng. | 1 |
| 2024 | An Explainable Automated Model for Measuring Software Engineer ContributionabstractSoftware engineers play an important role throughout the software development life-cycle, particularly in industry emphasizing quality assurance and timely delivery. Contribution measurement provides proper incentives to software engineers that motivate them to continuously improve the quality and efficiency of their work. However, existing research tends to ignore contribution measurement for software engineers in practice, relying heavily on peer review and lacking objectivity and transparency. Specifically, these studies still have two weaknesses. First, a few studies explore which metrics can be useful for contribution measurement in practice. Second, managers measure the contribution of software engineers based on their experience and lack of explainable automated tools to assist them. Yue Li 0047, He Zhang 0001, Yuzhe Jin, Liming Dong 0001, Lanxin Yang, David Lo 0001, Dong Shao |
ASE | 3 |
| 2016 | CO-GPS: Energy Efficient GPS Sensing with Cloud OffloadingabstractLocation is a fundamental service for mobile computing. Typical GPS receivers, although widely available for navigation purposes, may consume too much energy to be useful for many applications. Observing that in many sensing scenarios, the location information can be post-processed when the data is uploaded to a server, we design a cloud-offloaded GPS (CO-GPS) solution that allows a sensing device to aggressively duty-cycle its GPS receiver and log just enough raw GPS signal for post-processing. Leveraging publicly available information such as GNSS satellite ephemeris and an Earth elevation database, a cloud service can derive good quality GPS locations from a few milliseconds of raw data. Using our design of a portable sensing device platform called CLEON, we evaluate the accuracy and efficiency of the solution. Compared to more than 30 seconds of heavy signal processing on standalone GPS receivers, we can achieve three orders of magnitude lower energy consumption per location tagging. Jie Liu 0001, Bodhi Priyantha, Ted Hart, Yuzhe Jin, Woo Suk Lee, Vijay Raghunathan, Heitor S. Ramos, Qiang Wang 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2014 | Energy efficient GPS acquisition with sparse-gps
Prasant Misra, Wen Hu 0001, Yuzhe Jin, Jie Liu 0001, Amanda Souza de Paula, Niklas Wirström, Thiemo Voigt |
IPSN | 3 |
| 2014 | COIN-GPS: indoor localization from direct GPS receivingabstractDue to poor signal strength, multipath effects, and limited on-device computation power, common GPS receivers do not work indoors. This work addresses these challenges by using a steerable, high-gain directional antenna as the front-end of a GPS receiver along with a robust signal processing step and a novel location estimation technique to achieve direct GPS-based indoor localization. By leveraging the computing power of the cloud, we accommodate longer signals for acquisition, and remove the requirement of decoding timestamps or ephemeris data from GPS signals. We have tested our system in 31 randomly chosen spots inside five single-story, indoor environments such as stores, warehouses and shopping centers. Our experiments show that the system is capable of obtaining location fixes from 20 of these spots with a median error of less than 10 m, where all normal GPS receivers fail. Shahriar Nirjon, Jie Liu 0001, Gerald DeJean, Bodhi Priyantha, Yuzhe Jin, Ted Hart |
MobiSys | 5 |
| 2014 | Entity linking at the tail: sparse signals, unknown entities, and phrase modelsabstractWeb search is seeing a paradigm shift from keyword based search to an entity-centric organization of web data. To support web search with this deeper level of understanding, a web-scale entity linking system must have 3 key properties: First, its feature extraction must be robust to the diversity of web documents and their varied writing styles and content structures. Second, it must maintain high-precision linking for "tail" (unpopular) entities that is robust to the existence of confounding entities outside of the knowledge base and entity profiles with minimal information. Finally, the system must represent large-scale knowledge bases with a scalable and powerful feature representation. We have built and deployed a web-scale unsupervised entity linking system for a commercial search engine that addresses these requirements by combining new developments in sparse signal recovery to identify the most discriminative features from noisy, free-text web documents; explicit modeling of out-of-knowledge-base entities to improve precision at the tail; and the development of a new phrase-unigram language model to efficiently capture high-order dependencies in lexical features. Using a knowledge base of 100M unique people from a popular social networking site, we present experimental results in the challenging domain of people-linking at the tail, where most entities have limited web presence. Our experimental results show that this system substantially improves on the precision-recall tradeoff over baseline methods, achieving precision over 95% with recall over 60%. Yuzhe Jin, Emre Kiciman, Kuansan Wang, Ricky Loynd |
WSDM | 1 |
| 2013 | Sparse lexical representation for semantic entity resolutionabstractThis paper addresses the problem of semantic entity resolution (SER), which aims to determine whether some or none of the entities in a knowledge base is mentioned in a given web document. The lexical features, e.g., words and phrases, which are critical to the resolution of the semantic entities are typically of a small amount compared to all lexical features in the web document, and therefore can be modeled as sparse signals. Two techniques leveraging the principles of sparse signal recovery are proposed to identify the sparse, salient lexical features: one technique, based on the Lasso algorithm with the l2-norm distance metric, attempts to recover all the salient lexical features at once; the other technique, namely Posterior Probability Pursuit (PPP), sequentially identifies salient features one after one using the negative log posterior probability as the distance metric. Using a knowledge base consisting of about 100 million entities, we show that the proposed techniques exploiting the sparsity nature underlying SER deliver substantial performance improvement over baseline methods without sparsity consideration, demonstrating the potentials of sparse signal techniques in entity-centric web information processing. Yuzhe Jin, Kuansan Wang, Emre Kiciman |
ICASSP | 1 |
| 2013 | SparseGPS: energy efficient GPS acquisition via sparse approximationabstractThe global positioning system (GPS) system is a dominant wireless technology that enables reliable location sensing for a diverse range of outdoor mobile applications. Following rising demands for location sensing, low-cost GPS receivers are becoming widely available; but their energy demands are still too high to be useful for many of these applications. For energy efficient GPS sensing, the possibility of offloading a few milliseconds of raw signal samples and leveraging the greater processing power of the cloud for obtaining a position fix is being actively investigated. In an attempt to reduce the energy cost of this data offloading operation, we propose SparseGPS: a lightweight GPS acquisition mechanism based on sparse approximation. Prasant Misra, Wen Hu 0001, Yuzhe Jin, Jie Liu 0001, Niklas Wirström, Thiemo Voigt |
SenSys | 3 |
| 2013 | Support Recovery of Sparse Signals in the Presence of Multiple Measurement VectorsabstractThis paper studies the performance limits in the support recovery of sparse signals based on multiple measurement vectors (MMV). An information-theoretic analytical framework inspired by the connection to the single-input multiple-output multiple-access channel communication is established to reveal the performance limits in the support recovery of sparse signals with fixed number of nonzero entries. Sharp sufficient and necessary conditions for asymptotically successful support recovery are derived in terms of the number of measurements per vector, the number of nonzero rows, the measurement noise level, and the number of measurement vectors. Through the interpretations of the results, the benefit of having MMV for sparse signal recovery is illustrated, thus providing a theoretical foundation to the performance improvement enabled by MMV as observed in many existing simulation results. In particular, it is shown that the structure (rank) of the matrix formed by the nonzero entries plays an important role in the performance limits of support recovery. Yuzhe Jin, Bhaskar D. Rao |
IEEE Trans. Inf. Theory | 1 |
| 2011 | MultiPass lasso algorithms for sparse signal recoveryabstractWe develop the MultiPass Lasso (MPL) algorithm for sparse signal recovery. MPL applies the Lasso algorithm in a novel, sequential manner and has the following important attributes. First, MPL improves the estimation of the support of the sparse signal by combining high quality estimates of its partial supports which are sequentially recovered via the Lasso algorithm in each iteration/pass. Second, the algorithm is capable of exploiting the dynamic range in the nonzero magnitudes. Preliminary theoretic analysis shows the potential performance improvement enabled by MPL over Lasso. In addition, we propose the Reweighted MultiPass Lasso algorithm which substitutes Lasso with MPL in each iteration of Reweighted ℓ1Minimization. Experimental results favorably support the advantages of the proposed algorithms in both reconstruction accuracy and computational efficiency, thereby supporting the potential of the MultiPass framework for algorithmic development. Yuzhe Jin, Bhaskar D. Rao |
ISIT | 1 |
| 2011 | Limits on Support Recovery of Sparse Signals via Multiple-Access Communication TechniquesabstractIn this paper, we consider the problem of exact support recovery of sparse signals via noisy linear measurements. The main focus is finding the sufficient and necessary condition on the number of measurements for support recovery to be reliable. By drawing an analogy between the problem of support recovery and the problem of channel coding over the Gaussian multiple-access channel (MAC), and exploiting mathematical tools developed for the latter problem, we obtain an information-theoretic framework for analyzing the performance limits of support recovery. Specifically, when the number of nonzero entries of the sparse signal is held fixed, the exact asymptotics on the number of measurements sufficient and necessary for support recovery is characterized. In addition, we show that the proposed methodology can deal with a variety of models of sparse signal recovery, hence demonstrating its potential as an effective analytical tool. Yuzhe Jin, Young-Han Kim 0001, Bhaskar D. Rao |
IEEE Trans. Inf. Theory | 1 |
| 2010 | Algorithms for robust linear regression by exploiting the connection to sparse signal recoveryabstractIn this paper, we develop algorithms for robust linear regression by leveraging the connection between the problems of robust regression and sparse signal recovery. We explicitly model the measurement noise as a combination of two terms; the first term accounts for regular measurement noise modeled as zero mean Gaussian noise, and the second term captures the impact of outliers. The fact that the latter outlier component could indeed be a sparse vector provides the opportunity to leverage sparse signal reconstruction methods to solve the problem of robust regression. Maximum a posteriori (MAP) based and empirical Bayesian inference based algorithms are developed for this purpose. Experimental studies on simulated and real data sets are presented to demonstrate the effectiveness of the proposed algorithms. Yuzhe Jin, Bhaskar D. Rao |
ICASSP | 1 |
| 2010 | Performance tradeoffs for exact support recovery of sparse signalsabstractWe study the tradeoffs between the number of measurements, the signal sparsity level, and the measurement noise level for exact support recovery of sparse signals via random noisy measurements. By drawing analogy between exact support recovery and communication over the Gaussian multiple access channel, and exploiting mathematical tools developed for the latter problem, we derive sharp asymptotic sufficient and necessary conditions for exact support recovery. Specifically, when the number of nonzero entries is held fixed, the exact asymptotics on the number of measurements for support recovery is developed. When the number of nonzero entries increases in certain manners, we obtain sufficient conditions tighter than existing results. The proposed information theoretic framework for analyzing the performance of support recovery is further demonstrated to be capable of dealing with a variety of sparse signal recovery models. Yuzhe Jin, Young-Han Kim 0001, Bhaskar D. Rao |
ISIT | 1 |
| 2008 | Insights into the stable recovery of sparse solutions in overcomplete representations using network information theoryabstractIn this paper, we examine the problem of overcomplete representations and provide new insights into the problem of stable recovery of sparse solutions in noisy environments. We establish an important connection between the inverse problem that arises in overcomplete representations and wireless communication models in network information theory. We show that the stable recovery of a sparse solution with a single measurement vector (SMV) can be viewed as decoding competing users simultaneously transmitting messages through a Multiple Access Channel (MAC) at the same rate. With multiple measurement vectors (MMV), we relate the inverse problem to the wireless communication scenario with a Multiple-Input Multiple-Output (MIMO) channel. In each case, based on the connection established between the two domains, we leverage channel capacity results with outage analysis to shed light on the fundamental limits of any algorithm to stably recover sparse solutions in the presence of noise. Our results explicitly indicate the conditions on the key model parameters, e.g. degree of overcompleteness, degree of sparsity, and the signal-to-noise ratio, to guarantee the existence of asymptotically stable reconstruction of the sparse source. Yuzhe Jin, Bhaskar D. Rao |
ICASSP | 1 |
| 2008 | Performance limits of matching pursuit algorithmsabstractIn this paper, we examine the performance limits of the Orthogonal Matching Pursuit (OMP) algorithm, which has proven to be effective in solving for sparse solutions to inverse problem arising in overcomplete representations. To identify these limits, we exploit the connection between sparse solution problem and multiple access channel (MAC) in wireless communication domain. The forward selective nature of OMP helps it to be recognized as a successive interference cancellation (SIC) scheme that decodes non-zero entries one at a time in a specific order. We leverage this SIC decoding order and utilize the criterion for successful decoding to develop the information-theoretic performance limitation for OMP, which involves factors such as dictionary dimension, signal-to-noise-ratio, and importantly, the relative behavior of the non-zeros entries. Supported by computer simulations, our proposed criterion is demonstrated to be asymptotically effective in explaining the behavior of OMP. Yuzhe Jin, Bhaskar D. Rao |
ISIT | 1 |
| 2007 | Spectral Estimation of Voiced Speech using a Family of MVDR EstimatesabstractWe present a robust approach to modeling voiced speech using a family of minimum variance distortionless response (MVDR) spectral estimates. The method exploits the fact that for a fixed model order, for a sinusoidal signal in noise, the MVDR estimate at the sinusoidal frequency is approximately related to the sinusoidal and noise power in a simple linear manner with the coefficients being dependent on the model order. Modeling voiced speech as a sum of harmonic signals, we then use the aforementioned relationship along with a least squares approach to combine a family of MVDR estimates (MVDR estimates of different orders) and develop a robust approach for modeling voiced speech. Experimental results of spectral estimation of sinusoids, synthetic vowels, and actual speech signals at SNR of 0 dB and 5 dB using this approach indicate an increased resolution in the estimated MVDR spectra. The MFCC computed from the MVDR spectra using this approach are also used for speaker identification experiments on the TEMIT database at various SNR. The results indicate a reasonable improvement in recognition performance when compared to the MFCC and the fixed order MVDR-MFCC. Rajesh M. Hegde, Yuzhe Jin, Bhaskar D. Rao |
ICASSP (4) | 2 |