Institution profile

Guide Labs

Industry researchnorthamerica · us
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Prototype Language Models

Jul 01, 2026

This work addresses the challenge of efficiently tracing the influence of training samples on outputs in large language models. The authors propose PRISM, an architecture that performs prediction via sparse non-negative prototype mixtures, where each prototype is anchored by clustering objectives to coherent neighborhoods in the training data, thereby explicitly linking predictions to specific training instances. PRISM enables highly efficient attribution—approximately 500× faster than baseline methods—and supports behavior modification without fine-tuning. The approach further incorporates Hessian curvature localization and a linear prototype controller for calibration. Evaluated across models ranging from 130M to 1.6B parameters, PRISM incurs at most a 2.5-point accuracy drop relative to dense baselines on downstream tasks; prototype calibration recovers about 3 points of accuracy, and selective suppression of prototypes can eliminate targeted behaviors without degrading generation quality.

0 citationsRead paper

scCBGM: Interpretable Single-Cell Counterfactual Editing

Jun 05, 2026

This study addresses the challenge posed by the combinatorial explosion in single-cell perturbation experiments, which hinders comprehensive exploration of cellular phenotypic mechanisms. To overcome this limitation, the authors propose the single-cell Concept Bottleneck Generative Model (scCBGM), the first adaptation of concept bottleneck architectures to single-cell data. By incorporating decoder skip connections and a cross-covariance penalty, scCBGM achieves disentangled representations without dimensional constraints and extends naturally to a flow-matching framework for precise counterfactual generation and editing. The method demonstrates strong compositional generalization and counterfactual prediction capabilities across multiple real-world datasets, with efficacy validated through both cell-level synthetic benchmarks—featuring ground-truth counterfactual labels—and population-level experimental data.

0 citationsRead paper

The Attribution Contract: Feature Attribution for Generative Language Models

May 21, 2026

This work addresses the ambiguity in feature attribution for generative language models, which often stems from a lack of explicit semantic grounding. The paper introduces the “attribution contract” framework, systematically identifying that disputes in attribution arise from inconsistent implicit assumptions about explanations. By formally defining the explained output, attributable features, assumptions about the generative process, fixed conditions, and model scoring, the framework standardizes attribution practices. Through case studies on autoregressive and diffusion-based language models, it clarifies the validity boundaries of different attribution settings and underscores that attribution methods must be evaluated in alignment with their corresponding contracts. This approach prevents misleading interpretations and establishes a new paradigm for trustworthy explanations in generative models.

0 citationsRead paper
Recent publications

Latest Papers

Prototype Language Models

Jul 01, 2026

This work addresses the challenge of efficiently tracing the influence of training samples on outputs in large language models. The authors propose PRISM, an architecture that performs prediction via sparse non-negative prototype mixtures, where each prototype is anchored by clustering objectives to coherent neighborhoods in the training data, thereby explicitly linking predictions to specific training instances. PRISM enables highly efficient attribution—approximately 500× faster than baseline methods—and supports behavior modification without fine-tuning. The approach further incorporates Hessian curvature localization and a linear prototype controller for calibration. Evaluated across models ranging from 130M to 1.6B parameters, PRISM incurs at most a 2.5-point accuracy drop relative to dense baselines on downstream tasks; prototype calibration recovers about 3 points of accuracy, and selective suppression of prototypes can eliminate targeted behaviors without degrading generation quality.

0 citationsRead paper

scCBGM: Interpretable Single-Cell Counterfactual Editing

Jun 05, 2026

This study addresses the challenge posed by the combinatorial explosion in single-cell perturbation experiments, which hinders comprehensive exploration of cellular phenotypic mechanisms. To overcome this limitation, the authors propose the single-cell Concept Bottleneck Generative Model (scCBGM), the first adaptation of concept bottleneck architectures to single-cell data. By incorporating decoder skip connections and a cross-covariance penalty, scCBGM achieves disentangled representations without dimensional constraints and extends naturally to a flow-matching framework for precise counterfactual generation and editing. The method demonstrates strong compositional generalization and counterfactual prediction capabilities across multiple real-world datasets, with efficacy validated through both cell-level synthetic benchmarks—featuring ground-truth counterfactual labels—and population-level experimental data.

0 citationsRead paper

The Attribution Contract: Feature Attribution for Generative Language Models

May 21, 2026

This work addresses the ambiguity in feature attribution for generative language models, which often stems from a lack of explicit semantic grounding. The paper introduces the “attribution contract” framework, systematically identifying that disputes in attribution arise from inconsistent implicit assumptions about explanations. By formally defining the explained output, attributable features, assumptions about the generative process, fixed conditions, and model scoring, the framework standardizes attribution practices. Through case studies on autoregressive and diffusion-based language models, it clarifies the validity boundaries of different attribution settings and underscores that attribution methods must be evaluated in alignment with their corresponding contracts. This approach prevents misleading interpretations and establishes a new paradigm for trustworthy explanations in generative models.

0 citationsRead paper