Institution profile

University of Oregon

Academic institutionnorthamerica · us
Official website
Research library107linked papers
Opportunities0open roles
Selected work

Representative Papers

A Queueing Theoretic Perspective on Low-Latency LLM Inference with Variable Token Length

Jul 07, 2024International Symposium on Modeling and Optimization in Mobile, Ad-Hoc and Wireless Networks

Variable-length outputs in LLM interactive serving induce significant inference queuing latency due to output-token–dependent service times. Method: We propose a unified theoretical framework integrating M/G/1 and batch-service queueing models, the first to treat output token count as a stochastic service time. We jointly optimize the max-token limit and batch scheduling policies—fixed, dynamic, and elastic—to characterize their distinct latency behaviors under output-length uncertainty. Contribution/Results: Our analysis reveals the dominant impact of long-tail requests on mean queuing delay. Event-driven simulations validate model accuracy (<5% error): setting max-token = 256 reduces mean queuing delay by 38%; under load fluctuations, elastic batching cuts delay by 22% versus fixed batching. The core contribution is establishing a quantitative relationship between output-length variability and system latency, enabling principled co-optimization of inference parameters.

12 citationsRead paper

Linked Barcode for Persistence Induced by Filtrations

Aug 04, 2026

Traditional persistence barcodes struggle to capture algebraic relationships among homology classes across different dimensions, limiting their discriminative power in identifying complex structures. This work proposes a novel framework called “linked barcodes,” which explicitly constructs dynamic links between adjacent-dimensional bars by tracking the (p+1)-dimensional chains that render p-dimensional cycles into boundaries and monitoring their evolution within filtered complexes. By incorporating a reference filtration to ensure representation stability, the method uniquely integrates cross-dimensional homological relationships into persistent homology descriptors. Empirical evaluations demonstrate that this approach significantly outperforms standard persistent homology techniques in tasks such as graph isomorphism detection and temporal network link prediction.

0 citationsRead paper

Exemplars in Disguise: Pure Exemplar Models Mimic Abstraction-First Learning

Aug 01, 2026

This study challenges the prevailing view that large language models acquire abstract knowledge before learning from individual instances. By constructing a purely memory-based generative model devoid of explicit abstract representations, and employing sensitivity analysis, input distribution modeling, and temporal evaluation, the authors demonstrate that current evaluation paradigms may produce misleading evidence of “abstraction-first” learning. They show that a model’s sensitivity to individual samples and the statistical properties of their input distribution can create the illusion of abstract reasoning—even when the model relies solely on memorized instances. The findings suggest that instance-specific and abstract knowledge may be inherently entangled within distributed representations, thereby undermining the empirical basis and theoretical assumptions underlying the abstraction-priority hypothesis.

0 citationsRead paper

Co-Evolving Graph and Text Memory for Training-Free Multi-Hop Question Answering

Jul 25, 2026

This work addresses the challenge of effectively integrating knowledge graphs and textual evidence for multi-hop question answering. To this end, it proposes a training-free, open architecture that dynamically coordinates structured relational knowledge and unstructured contextual information through a synchronous bidirectional graph-text working memory mechanism. The core innovation lies in a co-evolutionary process between graph and text memories, which enables continuous alignment and mutual enhancement during both retrieval and generation stages. This is achieved via synchronized recurrent integration, relation triple extraction, and graph fact injection strategies. Evaluated on six mainstream multi-hop QA benchmarks, the method substantially outperforms existing training-free baselines and achieves performance comparable to larger-scale or trainable systems.

0 citationsRead paper

Updating zigzag representatives efficiently

Jul 17, 2026

This work addresses the inefficiency in updating persistence representatives under dynamic zigzag filtrations, where changes such as insertion or deletion of simplices alter adjacency relations and hinder computational performance. To overcome this challenge, the paper introduces a novel algorithm for extracting and updating zigzag persistence representatives based on the R = DV decomposition commonly used in non-zigzag settings. This approach achieves, for the first time, efficient maintenance of zigzag representatives amid evolving adjacency structures, employing an update strategy with quadratic time complexity. The proposed method significantly narrows the computational gap between zigzag and non-zigzag persistence computations, thereby substantially enhancing the efficiency of persistent homology calculations over dynamic filtrations.

0 citationsRead paper
Recent publications

Latest Papers

Linked Barcode for Persistence Induced by Filtrations

Aug 04, 2026

Traditional persistence barcodes struggle to capture algebraic relationships among homology classes across different dimensions, limiting their discriminative power in identifying complex structures. This work proposes a novel framework called “linked barcodes,” which explicitly constructs dynamic links between adjacent-dimensional bars by tracking the (p+1)-dimensional chains that render p-dimensional cycles into boundaries and monitoring their evolution within filtered complexes. By incorporating a reference filtration to ensure representation stability, the method uniquely integrates cross-dimensional homological relationships into persistent homology descriptors. Empirical evaluations demonstrate that this approach significantly outperforms standard persistent homology techniques in tasks such as graph isomorphism detection and temporal network link prediction.

0 citationsRead paper

Exemplars in Disguise: Pure Exemplar Models Mimic Abstraction-First Learning

Aug 01, 2026

This study challenges the prevailing view that large language models acquire abstract knowledge before learning from individual instances. By constructing a purely memory-based generative model devoid of explicit abstract representations, and employing sensitivity analysis, input distribution modeling, and temporal evaluation, the authors demonstrate that current evaluation paradigms may produce misleading evidence of “abstraction-first” learning. They show that a model’s sensitivity to individual samples and the statistical properties of their input distribution can create the illusion of abstract reasoning—even when the model relies solely on memorized instances. The findings suggest that instance-specific and abstract knowledge may be inherently entangled within distributed representations, thereby undermining the empirical basis and theoretical assumptions underlying the abstraction-priority hypothesis.

0 citationsRead paper

Co-Evolving Graph and Text Memory for Training-Free Multi-Hop Question Answering

Jul 25, 2026

This work addresses the challenge of effectively integrating knowledge graphs and textual evidence for multi-hop question answering. To this end, it proposes a training-free, open architecture that dynamically coordinates structured relational knowledge and unstructured contextual information through a synchronous bidirectional graph-text working memory mechanism. The core innovation lies in a co-evolutionary process between graph and text memories, which enables continuous alignment and mutual enhancement during both retrieval and generation stages. This is achieved via synchronized recurrent integration, relation triple extraction, and graph fact injection strategies. Evaluated on six mainstream multi-hop QA benchmarks, the method substantially outperforms existing training-free baselines and achieves performance comparable to larger-scale or trainable systems.

0 citationsRead paper

Updating zigzag representatives efficiently

Jul 17, 2026

This work addresses the inefficiency in updating persistence representatives under dynamic zigzag filtrations, where changes such as insertion or deletion of simplices alter adjacency relations and hinder computational performance. To overcome this challenge, the paper introduces a novel algorithm for extracting and updating zigzag persistence representatives based on the R = DV decomposition commonly used in non-zigzag settings. This approach achieves, for the first time, efficient maintenance of zigzag representatives amid evolving adjacency structures, employing an update strategy with quadratic time complexity. The proposed method significantly narrows the computational gap between zigzag and non-zigzag persistence computations, thereby substantially enhancing the efficiency of persistent homology calculations over dynamic filtrations.

0 citationsRead paper

Quantum Sampling Architecture for Protein Structure Reconstruction on Utility-Scale Hardware

Jul 07, 2026

This work addresses the challenge of predicting short peptide conformations within protein binding pockets—a task traditionally hindered by computationally expensive physics-based conformational searches. The authors propose QSAD, a novel framework that formulates peptide structure prediction as a non-iterative quantum sampling problem over amino acid–level Hamiltonians. Implemented on IBM Heron R2 via a quantum–classical hybrid architecture, QSAD eliminates conventional iterative optimization and instead employs noise-tolerant, coarse-grained quantum sampling. This approach significantly enhances prediction accuracy and robustness while enabling approximate reconstruction of the protein energy landscape. Evaluated on 101 peptide complexes, QSAD outperforms current AI and quantum baselines by 27–71% in accuracy, achieves the lowest variance, tolerates hardware-level noise at 3–5 times typical error rates, and accelerates quantum execution by 27× compared to VQE.

0 citationsRead paper