Institution profile

Smallest.ai

Industry research
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech

Jun 18, 2026

This study addresses the limited understanding of how natural language instructions modulate acoustic outputs in current stylized text-to-speech (TTS) systems, which hinders model controllability and failure attribution. For the first time, the diffusion attention attribution method (DAAM) is introduced to the speech generation domain to perform cross-attention attribution analysis across 25 network layers and 24 ODE steps of the CapSpeech-TTS model, enabling fine-grained visualization and quantification of the influence of style-descriptive words. The findings reveal that style words exert a global regulatory effect, with their attention intensity significantly correlated with fundamental frequency and energy. This influence is most pronounced in deeper network layers—particularly layer 17, which exhibits the strongest selectivity—and during early ODE integration steps.

0 citationsRead paper

Rewriting TTS Inference Economics: Lightning V2 on Tenstorrent Achieves 4x Lower Cost Than NVIDIA L40S

Mar 24, 2026

This work addresses the challenge of achieving both computational efficiency and high audio fidelity in low-precision text-to-speech (TTS) inference, which traditionally suffers from audible distortions and spectral artifacts. The authors propose a precision-aware architecture co-designed with hardware-software optimizations, enabling the first production-grade TTS system on the Tenstorrent platform to deploy 95% low-fidelity computations and 80% BlockFloat8 operations without perceptible quality degradation. By integrating BlockFloat8 quantization, an on-chip network (NoC), distributed SRAM, and a deterministic execution model, the approach substantially reduces memory traffic and computational overhead. Compared to an NVIDIA L40S GPU, the solution achieves comparable throughput at approximately one-fourth the accelerator cost while preserving high-fidelity audio output, thereby redefining the economics of real-time speech synthesis.

0 citationsRead paper
Recent publications

Latest Papers

How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech

Jun 18, 2026

This study addresses the limited understanding of how natural language instructions modulate acoustic outputs in current stylized text-to-speech (TTS) systems, which hinders model controllability and failure attribution. For the first time, the diffusion attention attribution method (DAAM) is introduced to the speech generation domain to perform cross-attention attribution analysis across 25 network layers and 24 ODE steps of the CapSpeech-TTS model, enabling fine-grained visualization and quantification of the influence of style-descriptive words. The findings reveal that style words exert a global regulatory effect, with their attention intensity significantly correlated with fundamental frequency and energy. This influence is most pronounced in deeper network layers—particularly layer 17, which exhibits the strongest selectivity—and during early ODE integration steps.

0 citationsRead paper

Rewriting TTS Inference Economics: Lightning V2 on Tenstorrent Achieves 4x Lower Cost Than NVIDIA L40S

Mar 24, 2026

This work addresses the challenge of achieving both computational efficiency and high audio fidelity in low-precision text-to-speech (TTS) inference, which traditionally suffers from audible distortions and spectral artifacts. The authors propose a precision-aware architecture co-designed with hardware-software optimizations, enabling the first production-grade TTS system on the Tenstorrent platform to deploy 95% low-fidelity computations and 80% BlockFloat8 operations without perceptible quality degradation. By integrating BlockFloat8 quantization, an on-chip network (NoC), distributed SRAM, and a deterministic execution model, the approach substantially reduces memory traffic and computational overhead. Compared to an NVIDIA L40S GPU, the solution achieves comparable throughput at approximately one-fourth the accelerator cost while preserving high-fidelity audio output, thereby redefining the economics of real-time speech synthesis.

0 citationsRead paper