bland-altman analysis

A method for assessing agreement between two measurement techniques by plotting differences versus means and computing limits of agreement to validate and quantify how new measurements compare to a reference standard.

bland-altmananalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Aug 01, 2026Aug 01, 2026
Career
Value
No comparison yet
$200K/year
Aug 01, 2026Aug 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

A statistical note on extending Christensen's limits of agreement with the mean

Aug 22, 2025
HS
Heidi Søgaard Christensen
🏛️ Aalborg University | Akershus University Hospital

This study addresses the neglect of subject–observer interaction effects in assessing inter-observer agreement for continuous measurements. We extend Christensen’s mean-based Limits of Agreement (LOAM) to a two-factor random-effects model incorporating interaction terms. By decomposing variance components, we rigorously distinguish between repeatability LOAM (within-observer) and reproducibility LOAM (between-observers), and derive their asymptotic confidence intervals, sample size formulas, and statistical tests for comparing LOAMs across measurement systems. The proposed framework unifies variance component estimation, LOAM inference, and hypothesis testing, thereby enhancing both the precision and interpretability of measurement system analysis. It provides a theoretically rigorous yet practically implementable tool for evaluating measurement consistency in clinical and biomedical research.

Developing reproducibility and repeatability metrics with confidence intervalsExtending LOAM to include subject-observer interaction effectsSeparating measurement error from systematic observer variation

This study addresses the critical need in clinical practice to assess whether new and existing measurement methods are interchangeable, which hinges on determining whether their results are clinically indistinguishable. To overcome the restrictive assumptions of current approaches—such as specific data distributions, homoscedasticity, and linear bias—the authors propose a more flexible inferential framework based on the Probability of Agreement (PoA). This framework integrates probabilistic modeling, statistical inference, and Monte Carlo simulation, thereby accommodating a broader range of real-world scenarios. The method is successfully demonstrated in a case study comparing tPSA measurement techniques and validated through extensive simulations, which confirm its robustness and superior performance. These advances substantially enhance the practical utility and generalizability of PoA-based interchangeability assessment.

clinical equivalenceflexible inferencemeasurement agreement

A new coefficient to measure agreement between continuous variables

Jul 10, 2025
RV
Ronny Vallejos
🏛️ Universidad Técnica Federico Santa María

In clinical studies, assessing agreement between two continuous measurement methods applied to the same subjects is a common yet challenging task. This paper proposes ρ₁, a novel agreement coefficient based on the L₁ distance, which requires no tuning parameters and exhibits strong robustness against outliers. Under bivariate normal and elliptically symmetric distributions, we rigorously derive its theoretical properties—including consistency, asymptotic normality, and invariance—and establish a complete statistical inference framework, supporting both confidence interval estimation and hypothesis testing. Extensive numerical experiments demonstrate that ρ₁ maintains stable performance across diverse distributions and contamination scenarios, consistently outperforming the classical Lin’s concordance correlation coefficient. By combining simplicity, robustness, and interpretability, ρ₁ provides a principled tool for agreement assessment in clinical comparisons, spatial analysis, and other applied settings requiring reliable quantification of method concordance.

Introduce robust L1-based coefficient rho1 for outlier resistanceMeasure agreement between continuous variables in clinical studiesValidate rho1 for bivariate normal and elliptical distributions

On the handling of method failure in comparison studies

Aug 21, 2024
MW
Milena Wunsch
🏛️ LMU Munich | Munich Center for Machine Learning | Department of Statistics | MRC Clinical Trials Unit | UCL

In methodological comparative studies, algorithmic failures—such as non-convergence or absence of output—preclude performance evaluation, yet existing literature lacks standardized guidelines for handling such failures, often overlooking or misapplying failure mitigation strategies. Method: We systematically analyze failure causes and risks of improper handling, critically examine prevalent censoring and imputation strategies for their statistical biases, and propose the principle of “context-adapted failure fallback,” establishing a framework grounded in empirically feasible fallback mechanisms. Through statistical modeling, failure root-cause diagnosis, and cross-domain empirical analysis, we identify widespread deficiencies in published studies’ failure handling practices. Contribution/Results: Two representative case studies demonstrate that inappropriate failure handling significantly distorts method rankings and undermines conclusion validity. Our work bridges critical theoretical and practical gaps in the principled treatment of algorithmic failures in empirical methodology research.

Addressing method failure handling in comparison studiesProviding guidance on proper failure interpretation and reportingRecommending realistic fallback strategies for method failures

Consistency is Key: Disentangling Label Variation in Natural Language Processing with Intra-Annotator Agreement

Jan 25, 2023
GA
Gavin Abercrombie
🏛️ Heriot-Watt University | Alana AI | Bocconi University

NLP data quality assessment has long relied on inter-annotator agreement, overlooking intra-annotator consistency—the temporal stability of individual annotators’ judgments. This neglect challenges the implicit “gold label as ground truth” assumption. Method: We conduct exploratory repeated annotation experiments across major NLP datasets and quantify intra-annotator agreement using Cohen’s and Fleiss’ Kappa, complemented by qualitative perceptual analysis. Contribution/Results: We demonstrate that mainstream NLP datasets routinely omit intra-annotator consistency reporting; moreover, individual annotators exhibit significant temporal variability in labeling identical texts. We identify and disentangle the dual influence of textual ambiguity and subjectivity on annotation stability. Our work establishes intra-annotator agreement as a foundational data quality metric, providing both a methodological framework and concrete guidelines for constructing more robust, reproducible NLP datasets.

Assessing annotator inconsistency across multiple NLP classification tasksInvestigating reasons for annotator disagreement through quality control measuresMeasuring intra-annotator agreement for label stability in NLP tasks

Latest Papers

What's happening recently
View more

Estimating the functional relationship between a continuous exposure and a binary outcome is challenging when covariates are measured with error. This study presents the first systematic evaluation of Simulation-Extrapolation, Regression Calibration, multiple imputation, and Bayesian correction methods, each coupled with flexible modeling techniques—including B-splines, P-splines, and fractional polynomials—within a multi-team, fully blinded, neutral simulation framework. By generating 155 distinct simulation scenarios and repeated samples, the research quantifies the bias and variance of each approach, revealing their relative strengths and limitations. The findings not only inform method selection under measurement error but also demonstrate the feasibility and value of this neutral comparative paradigm for rigorous methodological assessment.

covariate adjustmentexposure-outcome relationshipfunctional form

This study addresses the unreliable estimation of repeatability, between-laboratory, and reproducibility variance components under ISO 5725 standards when sample sizes are small or variance structures are extreme. To overcome this limitation, the authors propose a tailored Bootstrap resampling strategy adapted to a one-way random effects model. The approach refines point estimates by adjusting within-laboratory resampling and constructs confidence intervals via a two-stage resampling scheme integrated with bias-corrected and accelerated (BCa) techniques. Extensive simulations and validation using real data from ISO 5725-4 demonstrate that the proposed method substantially improves estimation accuracy and confidence interval coverage. It yields reliable, near-nominal or conservatively valid inferences for small- to moderate-sized experiments and clearly delineates optimal strategies across different practical scenarios.

bootstrapinterlaboratory precisionISO 5725

This study addresses the overreliance on inter-annotator agreement in current data annotation practices, which often overlooks annotation’s capacity to capture conceptual validity as a measurement act. Treating annotation as a measurement process, the work identifies five root causes of annotation issues—errors, ambiguity, impossibility, subjectivity, and annotator identity—and develops a measurement theory–based framework for diagnosing and improving annotation quality. Drawing on a synthesis of 132 literature sources and 10 semi-structured interviews, the research systematically defines target constructs, designs annotation instruments, implements labeling procedures, and evaluates both reliability and validity. The resulting framework equips annotation teams with evaluation methods that transcend mere agreement metrics, thereby substantially strengthening the foundational quality of AI training data.

annotation qualitydata annotationmeasurement

This study addresses the challenge of disentangling sources of inter-laboratory variability—specifically baseline offsets versus differences in sensitivity—in multi-laboratory assessments of linear dose–response relationships. To this end, the authors propose a precision evaluation framework based on linear mixed-effects models, integrating analysis of variance, F-tests, and ISO 5725 standards to define and estimate repeatability and between-laboratory variance components. Overall measurement precision is quantified via average dose-specific variance. Under a fully balanced design, the framework yields an exact decomposition of total sum of squares and closed-form ANOVA estimators, overcoming the limitation of conventional fixed-effects models that detect only the presence of differences without identifying their origin. The approach was successfully applied to bronchoalveolar lavage fluid data from a rat intratracheal instillation study involving nanomaterials, effectively distinguishing the sources of observed variability.

between-laboratory variancedose-response relationshipinterlaboratory studies

This study addresses the critical yet underexplored issue of how calibration and dichotomization thresholds in Qualitative Comparative Analysis (QCA) substantially influence analytical outcomes, while existing approaches lack systematic and efficient tools for sensitivity analysis. To bridge this gap, we introduce TSQCA, an R package that explicitly treats thresholds as analytical variables. TSQCA implements four sweep functions—otSweep, ctSweepS, ctSweepM, and dtSweep—to automate the exploration of multidimensional threshold combinations and their effects on QCA results. Built upon the CRAN QCA package for truth table construction and Boolean minimization, TSQCA employs an S3 object system to standardize output formats and supports automated generation of reproducible Markdown reports and visualizations. This framework significantly enhances the robustness, transparency, and reproducibility of QCA research.

calibration thresholdsdichotomizationQualitative Comparative Analysis

Hot Scholars

NL

Nils Lid Hjort

Professor of Mathematical Statistics, University of Oslo
Theoretical and applied statistics and probability theory