Institution profile

Home Team Science and Technology Agency

Industry researchasia · sg
Official website
Research library17linked papers
Opportunities0open roles
Selected work

Representative Papers

QCaption: Video Captioning and Q&A through Fusion of Large Multimodal Models

Jul 08, 2024Fusion

This work proposes a fully self-contained multimodal fusion architecture to address the limitations of insufficient semantic understanding and reliance on external services in video captioning and question-answering tasks. By integrating keyframe extraction, a large vision-language model (LVM), and a large language model (LLM), the framework enables end-to-end, efficient video understanding without dependence on external APIs, thereby supporting fully local deployment. Experimental results demonstrate significant performance gains, with up to 44.2% improvement in video captioning and 48.9% enhancement in video-based question answering, substantially advancing the system’s accuracy, practicality, and deployability in real-world scenarios.

1 citationsRead paper

Evaluating Prompt Scope and Demonstration Similarity in Local LLM Machine Translation

Jul 28, 2026

This study addresses the limited scope of current evaluations of large language models (LLMs) in machine translation, which often rely on a single prompt format and lack systematic analysis of multilingual target requests and example selection strategies. The work proposes a novel evaluation framework for local LLMs that treats prompt scope—monolingual versus language-family-level multilingual—and example retrieval strategy—random, lexical similarity, or embedding similarity—as key variables. Using the FLORES dataset under zero-shot and 5-shot settings, the authors evaluate LLaMA3.2-3B, Mistral, Qwen2.5-14B, and dedicated MT systems. Results show that dedicated MT systems consistently outperform LLMs; few-shot prompting improves Mistral and Qwen2.5 but degrades LLaMA3.2 performance; embedding-based retrieval slightly surpasses other strategies; and while language-family-level prompting is viable, smaller models are prone to structured output errors.

0 citationsRead paper

Transcoders for Investigating Deception in Language Models

Jul 16, 2026

This study addresses the critical security risks posed by deceptive behaviors in large language models and proposes an interpretable mechanism for their detection and intervention. For the first time, per-layer transcoders (PLTs) are applied to conduct circuit-level analysis of deception in large models. By constructing attribution graphs for Qwen3-4B, manipulating internal features, and tracing their downstream effects on model outputs, the work reveals that deceptive behavior is driven by specific internal mechanisms. The research successfully identifies a set of high-impact features associated with deception and compiles the first such feature dictionary. These findings demonstrate the effectiveness and potential of transcoders in behavioral monitoring and early detection of malicious intent within language models.

0 citationsRead paper

Efficiently Adapting Spoken Language Models for the Singaporean Context

Jul 10, 2026

This work addresses the challenge of efficiently adapting general-purpose speech-language models to Singapore’s multilingual, sociolinguistically sensitive contexts—such as those encountered by the Home Team—when access to original training data is unavailable. The authors propose an integrated approach combining LoRA-based fine-tuning, a novel multilingual question-answering dataset (HTD-multilingual-QA) designed to mitigate catastrophic forgetting, and an enhanced CoBa reweighting strategy extended for the first time to speech-based multitask learning. The resulting 5B-parameter model, HT-Moonstone, matches or surpasses the performance of models seven times its size across five multilingual speech tasks, achieving substantial gains in accent and gender recognition while incurring less than a 2% degradation in original speech question-answering capability.

0 citationsRead paper

Semantic Step Prediction: Multi-Step Latent Forecasting in LLM Reasoning Trajectories via Step Sampling

Apr 20, 2026

This work proposes Semantic Tube Prediction (STP) to enhance the predictability and geometric structure of hidden-state trajectories in large language models during multi-step reasoning. STP samples at semantic reasoning boundaries and applies geometric regularization to align trajectories with locally linear geodesics. The study reveals, for the first time, the critical influence of sampling positions on regularization efficacy, introduces a novel evaluation metric—multi-step hidden-state prediction mean squared error—and demonstrates that STP-generated trajectories are smooth curves rather than straight lines. Experiments on ProcessBench show that STP improves prediction accuracy by 168× over a frozen baseline; a 3-layer MLP trajectory predictor reduces error by 3–12× compared to linear extrapolation; and decoupling the language modeling loss doubles trajectory predictability.

0 citationsRead paper
Recent publications

Latest Papers

Evaluating Prompt Scope and Demonstration Similarity in Local LLM Machine Translation

Jul 28, 2026

This study addresses the limited scope of current evaluations of large language models (LLMs) in machine translation, which often rely on a single prompt format and lack systematic analysis of multilingual target requests and example selection strategies. The work proposes a novel evaluation framework for local LLMs that treats prompt scope—monolingual versus language-family-level multilingual—and example retrieval strategy—random, lexical similarity, or embedding similarity—as key variables. Using the FLORES dataset under zero-shot and 5-shot settings, the authors evaluate LLaMA3.2-3B, Mistral, Qwen2.5-14B, and dedicated MT systems. Results show that dedicated MT systems consistently outperform LLMs; few-shot prompting improves Mistral and Qwen2.5 but degrades LLaMA3.2 performance; embedding-based retrieval slightly surpasses other strategies; and while language-family-level prompting is viable, smaller models are prone to structured output errors.

0 citationsRead paper

Transcoders for Investigating Deception in Language Models

Jul 16, 2026

This study addresses the critical security risks posed by deceptive behaviors in large language models and proposes an interpretable mechanism for their detection and intervention. For the first time, per-layer transcoders (PLTs) are applied to conduct circuit-level analysis of deception in large models. By constructing attribution graphs for Qwen3-4B, manipulating internal features, and tracing their downstream effects on model outputs, the work reveals that deceptive behavior is driven by specific internal mechanisms. The research successfully identifies a set of high-impact features associated with deception and compiles the first such feature dictionary. These findings demonstrate the effectiveness and potential of transcoders in behavioral monitoring and early detection of malicious intent within language models.

0 citationsRead paper

Efficiently Adapting Spoken Language Models for the Singaporean Context

Jul 10, 2026

This work addresses the challenge of efficiently adapting general-purpose speech-language models to Singapore’s multilingual, sociolinguistically sensitive contexts—such as those encountered by the Home Team—when access to original training data is unavailable. The authors propose an integrated approach combining LoRA-based fine-tuning, a novel multilingual question-answering dataset (HTD-multilingual-QA) designed to mitigate catastrophic forgetting, and an enhanced CoBa reweighting strategy extended for the first time to speech-based multitask learning. The resulting 5B-parameter model, HT-Moonstone, matches or surpasses the performance of models seven times its size across five multilingual speech tasks, achieving substantial gains in accent and gender recognition while incurring less than a 2% degradation in original speech question-answering capability.

0 citationsRead paper

Semantic Step Prediction: Multi-Step Latent Forecasting in LLM Reasoning Trajectories via Step Sampling

Apr 20, 2026

This work proposes Semantic Tube Prediction (STP) to enhance the predictability and geometric structure of hidden-state trajectories in large language models during multi-step reasoning. STP samples at semantic reasoning boundaries and applies geometric regularization to align trajectories with locally linear geodesics. The study reveals, for the first time, the critical influence of sampling positions on regularization efficacy, introduces a novel evaluation metric—multi-step hidden-state prediction mean squared error—and demonstrates that STP-generated trajectories are smooth curves rather than straight lines. Experiments on ProcessBench show that STP improves prediction accuracy by 168× over a frozen baseline; a 3-layer MLP trajectory predictor reduces error by 3–12× compared to linear extrapolation; and decoupling the language modeling loss doubles trajectory predictability.

0 citationsRead paper

SMT-AD: a scalable quantum-inspired anomaly detection approach

Apr 06, 2026

This work addresses the limited efficiency and scalability of anomaly detection in high-dimensional data by proposing a quantum-inspired approach based on multi-resolution tensor superposition. The method integrates Fourier-assisted feature embedding with a low-rank matrix product operator architecture to enable efficient modeling: its parameter count grows linearly with feature dimensionality, supports high parallelizability, and automatically emphasizes salient features. Evaluated on standard benchmarks such as credit card transaction datasets, the proposed approach achieves detection performance comparable to or better than state-of-the-art baselines while substantially reducing model complexity, thereby offering a lightweight yet highly scalable solution for high-dimensional anomaly detection.

0 citationsRead paper