Institution profile

OKESTRO Co., Ltd.

Industry researchasia · kr
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Bridging Temporal and Textual Modalities: A Multimodal Framework for Automated Cloud Failure Root Cause Analysis

Jan 08, 2026arXiv.org

This work addresses the challenges of fusing heterogeneous modalities—such as time-series metrics and textual logs—and the inherent difficulty large language models face in processing continuous temporal data for root cause analysis in cloud infrastructure failures. To this end, the authors propose a multimodal diagnostic framework that aligns time-series performance indicators with the embedding space of pretrained language models through temporal semantic compression, a gated cross-attention alignment encoder, and a retrieval-augmented generation mechanism. This integration enables automated root cause localization informed by historical knowledge. Experimental evaluation across six cloud system benchmarks demonstrates that the proposed method achieves a diagnosis accuracy of 48.75%, significantly outperforming existing approaches, particularly in complex, multi-fault scenarios.

1 citationsRead paper

Quantifying Autoscaler Vulnerabilities: An Empirical Study of Resource Misallocation Induced by Cloud Infrastructure Faults

Jan 08, 2026arXiv.org

This study addresses how cloud infrastructure failures distort performance metrics, thereby misleading autoscaling systems and leading to resource misallocation, increased costs, or degraded service reliability. Through controlled simulations, the authors systematically evaluate the impact of four common failure types—including storage and network routing issues—on both vertical and horizontal scaling strategies across varying instance configurations and SLO thresholds. The work presents the first quantitative analysis of how such failures bias scaling decisions, revealing that horizontal scaling is particularly sensitive to transient anomalies. It further proposes design principles to distinguish genuine workload changes from failure-induced artifacts. Experimental results demonstrate that storage failures can incur up to $258 in additional monthly costs under horizontal scaling, while routing anomalies consistently cause resource under-provisioning, offering empirical foundations for building fault-tolerant autoscaling mechanisms.

0 citationsRead paper
Recent publications

Latest Papers

Bridging Temporal and Textual Modalities: A Multimodal Framework for Automated Cloud Failure Root Cause Analysis

Jan 08, 2026arXiv.org

This work addresses the challenges of fusing heterogeneous modalities—such as time-series metrics and textual logs—and the inherent difficulty large language models face in processing continuous temporal data for root cause analysis in cloud infrastructure failures. To this end, the authors propose a multimodal diagnostic framework that aligns time-series performance indicators with the embedding space of pretrained language models through temporal semantic compression, a gated cross-attention alignment encoder, and a retrieval-augmented generation mechanism. This integration enables automated root cause localization informed by historical knowledge. Experimental evaluation across six cloud system benchmarks demonstrates that the proposed method achieves a diagnosis accuracy of 48.75%, significantly outperforming existing approaches, particularly in complex, multi-fault scenarios.

1 citationsRead paper

Quantifying Autoscaler Vulnerabilities: An Empirical Study of Resource Misallocation Induced by Cloud Infrastructure Faults

Jan 08, 2026arXiv.org

This study addresses how cloud infrastructure failures distort performance metrics, thereby misleading autoscaling systems and leading to resource misallocation, increased costs, or degraded service reliability. Through controlled simulations, the authors systematically evaluate the impact of four common failure types—including storage and network routing issues—on both vertical and horizontal scaling strategies across varying instance configurations and SLO thresholds. The work presents the first quantitative analysis of how such failures bias scaling decisions, revealing that horizontal scaling is particularly sensitive to transient anomalies. It further proposes design principles to distinguish genuine workload changes from failure-induced artifacts. Experimental results demonstrate that storage failures can incur up to $258 in additional monthly costs under horizontal scaling, while routing anomalies consistently cause resource under-provisioning, offering empirical foundations for building fault-tolerant autoscaling mechanisms.

0 citationsRead paper