Chunxiao Zhou

dblp:02/5745 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
3since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Information extraction and text analysis · 90% 3D vision · 10%
Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 100%
Theoretical computer science
3 papers
Mathematical optimization · 34% Algorithms and data structures · 33% Combinatorics and discrete mathematics · 20%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%
Computer graphics and multimedia
2 papers
Computational photography and imaging · 53% Image and video processing · 47%

Topics — the 10 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computational social science and digital humanities
public administration
0.212016
Batch Model for Batched Timestamps Data Analysis with Application to the SSA Disability Program · KDD 2016
Mathematical optimization › least squares
constrained least squares
0.212016
Batch Model for Batched Timestamps Data Analysis with Application to the SSA Disability Program · KDD 2016
Algorithms and data structures › randomized algorithms › sampling
resampling algorithms
0.112012
Fast Resampling Weighted v-Statistics · NIPS 2012
Image and video processing › image decomposition › image separation
layer separation
0.112009
Gradient domain layer separation under independent motion · ICCV 2009
Information theory
hypothesis testing
0.112009
Efficient Moments-based Permutation Tests · NIPS 2009
Algorithms and data structures
permutation test
0.112009
Efficient Moments-based Permutation Tests · NIPS 2009
Performance modeling and evaluation
queueing models
0.112016
Batch Model for Batched Timestamps Data Analysis with Application to the SSA Disability Program · KDD 2016
Computer vision › 3D vision › depth estimation
dense depth estimation
0.012003
Dynamic Depth Recovery from Unsynchronized Video Streams · CVPR (2) 2003
Computer vision › 3D vision
depth estimation
0.012003
Dynamic Depth Recovery from Unsynchronized Video Streams · CVPR (2) 2003
Computational photography and imaging
multi-perspective imaging
0.012003
Dynamic Depth Recovery from Unsynchronized Video Streams · CVPR (2) 2003

Methods — techniques the papers use, named apart from their topics

randomization model · 1.5distribution analysis · 1.5batch search algorithm · 0.8constrained least squares · 0.5constrained least-squares · 0.2resampling · 0.1orbit enumeration · 0.1recursive moment derivation · 0.1poisson equation · 0.1pearson distribution series approximation · 0.1mutual information · 0.1gradient domain optimization · 0.1temporal offset estimation · 0.1stereo matching · 0.1optical flow · 0.1
YearPublicationVenuePosition
2024 Estimating Agreement by Chance for Sequence Annotation
abstract
In the field of natural language processing, correction of performance assessment for chance agreement plays a crucial role in evaluating the reliability of annotations.However, there is a notable dearth of research focusing on chance correction for assessing the reliability of sequence annotation tasks, despite their widespread prevalence in the field.To address this gap, this paper introduces a novel model for generating random annotations, which serves as the foundation for estimating chance agreement in sequence annotation tasks.Utilizing the proposed randomization model and a related comparison approach, we successfully derive the analytical form of the distribution, enabling the computation of the probable location of each annotated text segment and subsequent chance agreement estimation.Through a combination simulation and corpus-based evaluation, we successfully assess its applicability and validate its accuracy and efficacy.
Diya Li, Carolyn P. Rosé, Ao Yuan, Chunxiao Zhou
ACL (1)4
2024 Distilling Multi-Scale Knowledge for Event Temporal Relation Extraction
abstract
Event Temporal Relation Extraction (ETRE) is paramount but challenging. Within a discourse, event pairs are situated at different distances or the so-called proximity bands. The temporal ordering communicated about event pairs where at more remote (i.e., "long'') or less remote (i.e., "short'') proximity bands are encoded differently. SOTA models have tended to perform well on events situated at either short or long proximity bands, but not both. Nonetheless, real-world, natural texts contain all types of temporal event-pairs. In this paper, we present MulCo : Distilling Mul ti-Scale Knowledge via Co ntrastive Learning, a knowledge co-distillation approach that shares knowledge across multiple event pair proximity bands to improve performance on all types of temporal datasets. Our experimental results show that MulCo successfully integrates linguistic cues pertaining to temporal reasoning across both short and long proximity bands and achieves new state-of-the-art results on several ETRE benchmark datasets.
Hao-Ren Yao, Luke Breitfeller, Aakanksha Naik, Chunxiao Zhou, Carolyn P. Rosé
CIKM4
2022 QA4IE: A Quality Assurance Tool for Information Extraction
abstract
Quality assurance (QA) is an essential though underdeveloped part of the data annotation process. Although QA is supported to some extent in existing annotation tools, comprehensive support for QA is not standardly provided. In this paper we contribute QA4IE, a comprehensive QA tool for information extraction, which can (1) detect potential problems in text annotations in a timely manner, (2) accurately assess the quality of annotations, (3) visually display and summarize annotation discrepancies among annotation team members, (4) provide a comprehensive statistics report, and (5) support viewing of annotated documents interactively. This paper offers a competitive analysis comparing QA4IE and other popular annotation tools and demonstrates its features, usage, and effectiveness through a case study. The Python code, documentation, and demonstration video are available publicly at https://github.com/CC-RMD-EpiBio/QA4IE.
Rafael Jiménez Silva, Kaushik Gedela, Alex Marr, Bart Desmet, Carolyn P. Rosé, Chunxiao Zhou
LREC6
2020 Curating Annotated Corpora for Functioning Information
Rafael Jiménez Silva, Chunxiao Zhou, Bart Desmet, Ayah Zirikly, Alex Marr, Maryanne Sacco, Pei-Shu Ho, Jonathan Camacho, Julia Porcino, Guy Divita, Elizabeth Rasch
AMIA2
2017 Inductive identification of functional status information and establishing a gold standard corpus: A case study on the Mobility domain
abstract
The importance of functional status information (FSI) has become increasingly evident in recent years [1, 2]. However, implementation, application, and normalization of FSI in health care and Electronic Health Records (EHRs) have been largely underexplored. The World Health Organization's International Classification of Functioning, Disability and Health (ICF) [3] is considered to be the international standard for describing and coding function and health states. Nevertheless, the ICF provides only a limited vocabulary for recognizing FSI descriptions, since its purpose is to organize concepts related to functioning rather than to provide a comprehensive terminology or a complete set of relations between concepts. While the free text portion of EHRs might provide a more complete picture of health status, treatment, and progress, current Natural Language Processing (NLP) methods largely focus on extracting medical conditions (e.g. diagnoses and symptoms, etc.). The absence of a standardized functional terminology and incompleteness of the ICF as a vocabulary source makes it challenging to build a NLP system to extract FSI from EHR free text. Our work takes the first step towards extraction of FSI from free text by systematically identifying the structure of FSI related to Mobility, a key domain of the ICF and an important domain in the determination of work disability. Our interdisciplinary research group inductively evaluated examples extracted from over 1,200 Physical Therapy (PT) notes from the Clinical Center of the National Institutes of Health (NIH). This extensive work resulted in a nested entity structure comprised of 2 entities, 3 sub-entities, 8 attributes, and 21 attribute values. Furthermore, we have manually curated the first gold standard corpus of 200 double-annotated and 50 triple-annotated PT notes. Our inter-annotator agreement (IAA) averages 97% F1-score on partial textual span matching and from 0.4 to 0.9 Siegel & Castellan's kappa on attribute value matching. Such a rich semantic corpus of Mobility FSI is valuable and a promising resource for future statistical learning. Our method is also adaptable to other domains of the ICF.
Thanh Thieu, Jonathan Camacho, Pei-Shu Ho, Julia Porcino, Lisa Nelson, Elizabeth Rasch, Chunxiao Zhou, Leighton Chan, Diane Brandt, Denis Newman-Griffis, Ao Yuan, Albert M. Lai
BIBM8
2016 Batch Model for Batched Timestamps Data Analysis with Application to the SSA Disability Program
abstract
The Office of Disability Adjudication and Review (ODAR) is responsible for holding hearings, issuing decisions, and reviewing appeals as part of the Social Security Administration's disability determining process. In order to control and process cases, the ODAR has established a Case Processing and Management System (CPMS) to record management information since December 2003. The CPMS provides a detailed case status history for each case. Due to the large number of appeal requests and limited resources, the number of pending claims at ODAR was over one million cases by March 31, 2015. Our National Institutes of Health (NIH) team collaborated with SSA and developed a Case Status Change Model (CSCM) project to meet the ODAR's urgent need of reducing backlogs and improve hearings and appeals process. One of the key issues in our CSCM project is to estimate the expected service time and its variation for each case status code. The challenge is that the systems recorded job departure times may not be the true job finished times. As the CPMS timestamps data of case status codes showed apparent batch patterns, we proposed a batch model and applied the constrained least squares method to estimate the mean service times and the variances. We also proposed a batch search algorithm to determine the optimal batch partition, as no batch partition was given in the real data. Simulation studies were conducted to evaluate the performance of the proposed methods. Finally, we applied the method to analyze a real CPMS data from ODAR/SSA.
Qingqi Yue, Ao Yuan, Xuan Che, Minh Huynh, Chunxiao Zhou
KDD5
2012 Fast Resampling Weighted v-Statistics
abstract
In this paper, a novel, computationally fast, and alternative algorithm for com- puting weighted v-statistics in resampling both univariate and multivariate data is proposed. To avoid any real resampling, we have linked this problem with finite group action and converted it into a problem of orbit enumeration. For further computational cost reduction, an efficient method is developed to list all orbits by their symmetry order and calculate all index function orbit sums and data function orbit sums recursively. The computational complexity analysis shows reduction in the computational cost from n! or nn level to low-order polynomial level.
Chunxiao Zhou, Jiseong Park, Yun Fu 0001
NIPS1
2009 Gradient domain layer separation under independent motion
abstract
Multi-exposure X-ray imaging can see through objects and separate different material into transparent layers. However, layer motion makes the separation task under-determined. Instead of aligning the non-rigid motion, we address the layer separation problem in gradient domain and propose an energy optimization framework to regularize it by explicitly enforcing independence constraint. It is shown that gradient domain allows more accurate and robust independence analysis between non-stationary signal using mutual information (MI) and hence achieves better separation. Furthermore, gradient fields contain sufficient information for full reconstruction of separated layers by solving the Poisson Equation. For efficient regularization of the gradient separation, energy terms based on the Taylor expansion of MI is further derived. Evaluation on both synthesized and real datasets proves the effectiveness of our algorithm and its robustness to complex tissue motion.
Yunqiang Chen, Ti-Chiun Chang, Chunxiao Zhou, Tong Fang
ICCV3
2009 Efficient Moments-based Permutation Tests
abstract
In this paper, we develop an efficient moments-based permutation test approach to improve the system’s efficiency by approximating the permutation distribution of the test statistic with Pearson distribution series. This approach involves the calculation of the first four moments of the permutation distribution. We propose a novel recursive method to derive these moments theoretically and analytically without any permutation. Experimental results using different test statistics are demonstrated using simulated data and real data. The proposed strategy takes advantage of nonparametric permutation tests and parametric Pearson distribution approximation to achieve both accuracy and efficiency.
Chunxiao Zhou, Huixia Judy Wang, Yongmei Michelle Wang
NIPS1
2008 3D face analysis for distinct features using statistical randomization
abstract
It is a fascinating yet challenging problem to accurately and efficiently localize regionally distinct features between face groups in multi-dimensional signal processing and analysis. Given a data with unknown distribution and small sample size, we propose a new statistical analysis framework using hybrid randomization (i.e., permutation) tests to improve the system's efficiency in identifying distinct features. The proposed method fits the nonparametric distribution of the test statistic with Pearson distribution series. We bypass the tedious online randomization via calculating the first four moments of the permutation distribution. This can reduce the computational complexity from O(n!) to O(n2) over traditional methods for the modified Hotelling's T2test statistics. Experiments on simulated data and 3D face analysis demonstrate the efficiency, accuracy and robustness of the proposed approach.
Chunxiao Zhou, Yuxiao Hu 0001, Yun Fu 0001, Huixia Judy Wang, Thomas S. Huang, Yongmei Michelle Wang
ICASSP1
2003 Dynamic Depth Recovery from Unsynchronized Video Streams
abstract
We propose an algorithm for estimating dense depth information of dynamic scenes from multiple video streams captured using unsynchronized stationary cameras. We solve this problem by first imposing two assumptions about the scene motion and the temporal offset between cameras. The motion of a scene is described using a local constant velocity model and the camera temporal offset is assumed to be constant within a short of period of time. Based on these models, geometric relations between the images of moving scene points, the scene depth, the scene motions, and the camera temporal offset are investigated and an estimation method is developed to compute the camera temporal offset. The three main steps of the proposed algorithm are: 1) the estimation of the temporal offset between cameras, 2) the synthesis of synchronized image pairs based on the estimated camera temporal offset and optical flow fields computed in each view, and 3) the stereo computation based on the synthesized synchronous image pairs. The proposed algorithm has been tested on both synthetic data and real image sequences. Promising quantitative and qualitative experimental results are demonstrated in the paper.
Chunxiao Zhou
CVPR (2)1