Score
Translating abstract constructs into concrete, measurable variables, annotation schemes, or protocol definitions so they can be reliably measured or labeled in data, including designing metrics, tiers, and procedures that reflect the target theoretical concept.
This study addresses the overreliance on inter-annotator agreement in current data annotation practices, which often overlooks annotation’s capacity to capture conceptual validity as a measurement act. Treating annotation as a measurement process, the work identifies five root causes of annotation issues—errors, ambiguity, impossibility, subjectivity, and annotator identity—and develops a measurement theory–based framework for diagnosing and improving annotation quality. Drawing on a synthesis of 132 literature sources and 10 semi-structured interviews, the research systematically defines target constructs, designs annotation instruments, implements labeling procedures, and evaluates both reliability and validity. The resulting framework equips annotation teams with evaluation methods that transcend mere agreement metrics, thereby substantially strengthening the foundational quality of AI training data.
Data scientists frequently lack systematic guidance when operationalizing ambiguous concepts (e.g., “writing authenticity,” “medical need”) into model-ready proxy target variables. To address this, we conducted semi-structured interviews with 15 data scientists across education and healthcare domains, followed by cross-domain thematic coding. We propose the “assemblage metrics” framework, identifying five core design criteria: validity, simplicity, predictiveness, portability, and resource efficiency. Our analysis reveals an iterative, problem-reconstruction–driven practice in which target variables are dynamically negotiated through trade-offs among these criteria. This work offers the first systematic characterization of such trade-offs in proxy target construction. It contributes a theoretically grounded framework and methodological tools for HCI, CSCW, and machine learning communities to support principled, transparent, and trustworthy predictive modeling—bridging conceptual abstraction with operationalizable measurement.
This study addresses the limited semantic transparency and poor comprehensibility of existing conceptual models, which stem from their reliance on low-level syntactic constructs to represent domain abstractions, thereby hindering effective system design and stakeholder communication. To overcome this, the paper proposes a language-agnostic abstract symbol engineering approach that identifies, formalizes, visualizes, and validates recurring syntactic configuration patterns, replacing them with high-level, semantically transparent abstract symbols. The method is instantiated as the DeCleaR extension to Dynamic Condition Response (DCR) graphs. Empirical evaluation demonstrates that DeCleaR significantly enhances perceived model quality, pragmatic quality, and user preference compared to standard DCR graphs.
Current approaches to automated program synthesis lack effective governance mechanisms to ensure the compliance of generated code. This work proposes Protocol-Driven Development (PDD), a model that treats machine-executable protocols as primary artifacts and delineates the space of valid implementations through structural, behavioral, and operational invariants. PDD mandates that every implementation be accompanied by a verifiable chain of compliance evidence. By integrating formal methods, property-based testing, policy-as-code, and software provenance techniques, PDD establishes a unified framework for protocol specification and verification. This framework enables trustworthy admission control over automatically synthesized code, guaranteeing that all adopted implementations strictly adhere to protocol constraints and are backed by complete, auditable proofs of compliance.
Scientists frequently record experimental metadata in spreadsheets, yet ensuring consistency and standards compliance remains challenging. This paper introduces a spreadsheet-native metadata governance paradigm: customized Excel/CSV templates embed HuBMAP standards; OWL/SKOS ontology-driven controlled vocabularies are integrated; and a web-based real-time semantic validation tool enables immediate, on-entry verification. The approach seamlessly incorporates semantic constraints into familiar spreadsheet workflows—requiring no platform switching or new system adoption. Deployed across the HuBMAP Consortium, it significantly improved multi-omics metadata compliance rates, increased data entry efficiency, and reduced error identification and correction time by over 70%. To our knowledge, this is the first work to deeply embed ontology-based constraints and real-time semantic validation directly within spreadsheet environments, establishing a scalable, practical paradigm for biomedical metadata standardization.
This work addresses the prevailing lack of systematic understanding of foundational formal theories in current AI compiler design, which hinders rigorous evaluation of the completeness and desirability of intermediate representations and compilation abstractions. For the first time, it systematically establishes precise correspondences between core mechanisms of MLIR—such as term rewriting systems, refinement calculi, and abstract interpretation—and classical formal theories. By grounding compiler abstractions in formal semantics, the paper clarifies the theoretical underpinnings of these constructs, articulates a precise notion of “design completeness,” and provides assessable criteria and guiding principles to navigate trade-offs between engineering pragmatism and theoretical ideals.
This study addresses the lack of systematic mechanisms in existing digital twin services to express and enforce data quality requirements—such as accuracy, completeness, and timeliness—at the model level during runtime, which undermines service reliability. To bridge this gap, the authors propose a contract-based approach to data quality management that formalizes a theory of data contracts, integrates it into the digital twin architecture, and introduces a domain-specific language (DSL) to support declarative specification and automated monitoring of these contracts. This work achieves, for the first time, a closed-loop integration of model-driven data contracts within digital twins, ensuring end-to-end data quality from the modeling phase through runtime execution. The proposed method significantly enhances the trustworthiness of downstream services, including simulation, what-if analysis, and machine learning–based prediction.
This study addresses the growing challenge posed by the widespread involvement of AI agents in software development, which undermines the long-standing assumption that development artifacts are exclusively produced by human professionals—an assumption underpinning traditional software metrics. The work systematically exposes how AI-generated traces compromise the foundational premises of established software measurement practices, thereby threatening the validity of prior empirical conclusions. To confront this issue, the authors propose an AI-augmented, systematic replication methodology that integrates modern data analytics with empirical software engineering techniques to rigorously re-evaluate key findings. The project advances a dynamic, reproducible, and sustainable measurement paradigm capable of adapting to evolving data ecosystems, offering a robust and timely framework for software metrics in the AI era.
Existing approaches to automatic formalization are largely confined to isolated statements and struggle to capture the intricate dependency structures among axioms, definitions, and lemmas within mathematical theories. This work introduces a novel paradigm—*theory-level automatic formalization*—which systematically advocates shifting from statement-level to integrated theory-level formalization. By constructing formalized mathematical libraries, modeling dependency graphs, and establishing mappings from natural to formal languages, the approach enables machine-verifiable translation of entire mathematical theories along with their internal structures. The paper delineates core challenges inherent to this direction, proposes three viable research pathways, and releases a comprehensive survey resource to catalyze further progress in the field.
Although large language models (LLMs) can achieve agreement with human annotators in text coding, their judgments may rely on superficial features unrelated to the underlying theoretical construct, thereby lacking construct validity. To address this issue, this work proposes a “fine-grained calibration” approach that decomposes theoretical constructs into clause-level components, validates each component against extractive evidence, and aggregates results according to explicit theoretical rules to assess whether LLMs genuinely measure the target construct. This method shifts the validation of construct validity from output consistency to process interpretability, enabling identification of errors stemming either from missing components or confusion with neighboring constructs. It establishes a transparent and interpretable paradigm for trustworthy measurement using LLMs in the social sciences.