Institution profile

Granica

Industry researchnorthamerica · us
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Generating from Discrete Distributions Using Diffusions: Insights from Random Constraint Satisfaction Problems

Mar 20, 2026

This work investigates efficient methods for uniformly sampling solutions to random k-SAT or k-XORSAT formulas to enhance the performance of discrete generative models on synthetic constraint satisfaction problem (CSP) benchmarks. The authors systematically compare continuous diffusion with masked discrete diffusion strategies and examine the impact of variable ordering on generation quality. Experimental results demonstrate that continuous diffusion models not only achieve the theoretically optimal accuracy but also significantly outperform existing discrete approaches. Moreover, specific variable orderings substantially improve generation fidelity without relying on conventional heuristic rules. These findings reveal non-intuitive influences of CSP theory on generative model behavior and offer a novel perspective for modeling discrete diffusion processes.

0 citationsRead paper

Train on Validation (ToV): Fast data selection with applications to fine-tuning

Sep 30, 2025

Data selection for fine-tuning under scarce target-distribution samples remains challenging. Method: This paper proposes a “validation-set-driven data selection” paradigm: it swaps the conventional roles of validation set and training pool—performing lightweight fine-tuning on the validation set and selecting the most discriminative samples from the training pool based on the magnitude of prediction shifts induced by fine-tuning. The method requires no additional annotations or gradient computations, ensuring both efficiency and theoretical interpretability. Results: Evaluated on instruction tuning and named entity recognition, it significantly reduces test log-loss on the target distribution, consistently outperforming existing SOTA methods on average while improving data utilization efficiency and fine-tuning performance. Its core innovation lies in the first use of the validation set as a proxy for fine-tuning and leveraging prediction shift as the selection criterion—enabling precise identification of high-information samples under few-shot settings.

0 citationsRead paper

Have ASkotch: A Neat Solution for Large-scale Kernel Ridge Regression

Jul 14, 2024

Large-scale kernel ridge regression (KRR) suffers from prohibitive computational and memory complexity, hindering its scalability to big-data regimes; while existing inducing-point approximations improve scalability, they incur substantial prediction accuracy loss. This paper introduces ASkotch—a novel, exact KRR solver achieving the first linear convergence guarantee independent of the condition number. Its core innovations integrate randomized preconditioning, accelerated gradient iteration, adaptive sampling via ridge leverage scores, and determinant point process theory—yielding an exact, bias-free, and provably convergent KRR framework. Evaluated on 23 large-scale regression and classification benchmarks across diverse domains, ASkotch consistently outperforms state-of-the-art methods in both accuracy and efficiency. It enables practical deployment of exact KRR in high-stakes applications such as computational chemistry and clinical prediction.

0 citationsRead paper
Recent publications

Latest Papers

Generating from Discrete Distributions Using Diffusions: Insights from Random Constraint Satisfaction Problems

Mar 20, 2026

This work investigates efficient methods for uniformly sampling solutions to random k-SAT or k-XORSAT formulas to enhance the performance of discrete generative models on synthetic constraint satisfaction problem (CSP) benchmarks. The authors systematically compare continuous diffusion with masked discrete diffusion strategies and examine the impact of variable ordering on generation quality. Experimental results demonstrate that continuous diffusion models not only achieve the theoretically optimal accuracy but also significantly outperform existing discrete approaches. Moreover, specific variable orderings substantially improve generation fidelity without relying on conventional heuristic rules. These findings reveal non-intuitive influences of CSP theory on generative model behavior and offer a novel perspective for modeling discrete diffusion processes.

0 citationsRead paper

Train on Validation (ToV): Fast data selection with applications to fine-tuning

Sep 30, 2025

Data selection for fine-tuning under scarce target-distribution samples remains challenging. Method: This paper proposes a “validation-set-driven data selection” paradigm: it swaps the conventional roles of validation set and training pool—performing lightweight fine-tuning on the validation set and selecting the most discriminative samples from the training pool based on the magnitude of prediction shifts induced by fine-tuning. The method requires no additional annotations or gradient computations, ensuring both efficiency and theoretical interpretability. Results: Evaluated on instruction tuning and named entity recognition, it significantly reduces test log-loss on the target distribution, consistently outperforming existing SOTA methods on average while improving data utilization efficiency and fine-tuning performance. Its core innovation lies in the first use of the validation set as a proxy for fine-tuning and leveraging prediction shift as the selection criterion—enabling precise identification of high-information samples under few-shot settings.

0 citationsRead paper

Have ASkotch: A Neat Solution for Large-scale Kernel Ridge Regression

Jul 14, 2024

Large-scale kernel ridge regression (KRR) suffers from prohibitive computational and memory complexity, hindering its scalability to big-data regimes; while existing inducing-point approximations improve scalability, they incur substantial prediction accuracy loss. This paper introduces ASkotch—a novel, exact KRR solver achieving the first linear convergence guarantee independent of the condition number. Its core innovations integrate randomized preconditioning, accelerated gradient iteration, adaptive sampling via ridge leverage scores, and determinant point process theory—yielding an exact, bias-free, and provably convergent KRR framework. Evaluated on 23 large-scale regression and classification benchmarks across diverse domains, ASkotch consistently outperforms state-of-the-art methods in both accuracy and efficiency. It enables practical deployment of exact KRR in high-stakes applications such as computational chemistry and clinical prediction.

0 citationsRead paper