Institution profile

Brave Software

Industry researchnorthamerica · us
Official website
Research library10linked papers
Opportunities0open roles
Selected work

Representative Papers

Towards Real-Time ECG and EMG Modeling on $μ$ NPUs

Apr 20, 2026

This work addresses the computational and memory bottlenecks of deploying high-performance Transformer models for electrocardiogram (ECG) and electromyogram (EMG) analysis on resource- and power-constrained micro neural processing units (μNPUs). To this end, the authors propose PhysioLite—a lightweight, hardware-aware model architecture and training framework that integrates learnable wavelet filter banks, CPU-offloaded positional encoding, μNPU-optimized network layers, and 8-bit quantization. This approach achieves state-of-the-art accuracy on ECG and EMG tasks while reducing model size to approximately 370 KB—less than 10% of the baseline—and demonstrates efficient real-time inference with low latency and power consumption. Notably, PhysioLite is the first to enable effective physiological signal modeling on real-world μNPU platforms such as the MAX78000 and HX6538 WE2.

0 citationsRead paper

SPILLage: Agentic Oversharing on the Web

Feb 13, 2026

This work addresses the unintended leakage of user privacy by large language model–driven web agents, which often disclose task-irrelevant personal information through their actions. We propose SPILLage, a framework that formally defines and quantifies the phenomenon of “natural agent oversharing,” introduces a two-dimensional taxonomy of privacy leakage—distinguishing content versus behavior and explicit versus implicit disclosures—and establishes a benchmark comprising 180 e-commerce tasks. Leveraging real-world website interaction traces annotated for task relevance, our analysis of 1,080 experimental runs reveals that behavioral leakage occurs five times more frequently than content leakage. Furthermore, we demonstrate that proactively removing irrelevant information via prompt engineering and input sanitization not only substantially reduces privacy exposure but also improves task success rates by up to 17.9%, uncovering a positive correlation between privacy preservation and task performance.

0 citationsRead paper

Privacy Practices of Browser Agents

Dec 08, 2025

This study identifies severe privacy risks in eight mainstream browser automation agents, stemming from component vulnerabilities, ineffective anti-tracking mechanisms, and unintended leakage of sensitive information during automated browsing. Method: We propose the first comprehensive privacy risk assessment framework for browser agents, evaluating five dimensions—data collection practices, tracking prevention efficacy, permission management, configuration security, and prompt robustness—across 15 rigorously defined metrics. Our empirical evaluation integrates static security auditing, dynamic behavioral analysis, and prompt-response testing across real-world websites. Contribution/Results: We uncover 30 previously undocumented privacy vulnerabilities (e.g., privacy-enhancing features disabled by default, automatic credential autofill). All findings were responsibly disclosed. To support reproducible research and evidence-based privacy governance, we will open-source our curated dataset and assessment toolkit. This work establishes a methodologically sound, empirically grounded foundation for evaluating and improving the privacy posture of browser automation agents.

0 citationsRead paper

Membership and Memorization in LLM Knowledge Distillation

Aug 09, 2025

This study is the first to systematically uncover dual privacy risks—membership inference and memorization leakage—in large language model (LLM) knowledge distillation (KD). Addressing privacy hazards introduced when teacher models are trained on sensitive data, we evaluate six mainstream KD methods across three model families (GPT-2, LLaMA-2, OPT) and seven NLP tasks, quantifying how distillation objectives, student data composition, and task types affect privacy leakage. We propose a novel “modular privacy analysis” framework, revealing substantial heterogeneity in privacy propagation across network blocks. Experimental results demonstrate that all existing LLM KD methods inherit and transmit teacher-model privacy risks—but to varying degrees—and critically, that membership inference and memorization leakage exhibit significant inconsistency, challenging the conventional assumption of their equivalence.

0 citationsRead paper
Recent publications

Latest Papers

Towards Real-Time ECG and EMG Modeling on $μ$ NPUs

Apr 20, 2026

This work addresses the computational and memory bottlenecks of deploying high-performance Transformer models for electrocardiogram (ECG) and electromyogram (EMG) analysis on resource- and power-constrained micro neural processing units (μNPUs). To this end, the authors propose PhysioLite—a lightweight, hardware-aware model architecture and training framework that integrates learnable wavelet filter banks, CPU-offloaded positional encoding, μNPU-optimized network layers, and 8-bit quantization. This approach achieves state-of-the-art accuracy on ECG and EMG tasks while reducing model size to approximately 370 KB—less than 10% of the baseline—and demonstrates efficient real-time inference with low latency and power consumption. Notably, PhysioLite is the first to enable effective physiological signal modeling on real-world μNPU platforms such as the MAX78000 and HX6538 WE2.

0 citationsRead paper

SPILLage: Agentic Oversharing on the Web

Feb 13, 2026

This work addresses the unintended leakage of user privacy by large language model–driven web agents, which often disclose task-irrelevant personal information through their actions. We propose SPILLage, a framework that formally defines and quantifies the phenomenon of “natural agent oversharing,” introduces a two-dimensional taxonomy of privacy leakage—distinguishing content versus behavior and explicit versus implicit disclosures—and establishes a benchmark comprising 180 e-commerce tasks. Leveraging real-world website interaction traces annotated for task relevance, our analysis of 1,080 experimental runs reveals that behavioral leakage occurs five times more frequently than content leakage. Furthermore, we demonstrate that proactively removing irrelevant information via prompt engineering and input sanitization not only substantially reduces privacy exposure but also improves task success rates by up to 17.9%, uncovering a positive correlation between privacy preservation and task performance.

0 citationsRead paper

Privacy Practices of Browser Agents

Dec 08, 2025

This study identifies severe privacy risks in eight mainstream browser automation agents, stemming from component vulnerabilities, ineffective anti-tracking mechanisms, and unintended leakage of sensitive information during automated browsing. Method: We propose the first comprehensive privacy risk assessment framework for browser agents, evaluating five dimensions—data collection practices, tracking prevention efficacy, permission management, configuration security, and prompt robustness—across 15 rigorously defined metrics. Our empirical evaluation integrates static security auditing, dynamic behavioral analysis, and prompt-response testing across real-world websites. Contribution/Results: We uncover 30 previously undocumented privacy vulnerabilities (e.g., privacy-enhancing features disabled by default, automatic credential autofill). All findings were responsibly disclosed. To support reproducible research and evidence-based privacy governance, we will open-source our curated dataset and assessment toolkit. This work establishes a methodologically sound, empirically grounded foundation for evaluating and improving the privacy posture of browser automation agents.

0 citationsRead paper

Membership and Memorization in LLM Knowledge Distillation

Aug 09, 2025

This study is the first to systematically uncover dual privacy risks—membership inference and memorization leakage—in large language model (LLM) knowledge distillation (KD). Addressing privacy hazards introduced when teacher models are trained on sensitive data, we evaluate six mainstream KD methods across three model families (GPT-2, LLaMA-2, OPT) and seven NLP tasks, quantifying how distillation objectives, student data composition, and task types affect privacy leakage. We propose a novel “modular privacy analysis” framework, revealing substantial heterogeneity in privacy propagation across network blocks. Experimental results demonstrate that all existing LLM KD methods inherit and transmit teacher-model privacy risks—but to varying degrees—and critically, that membership inference and memorization leakage exhibit significant inconsistency, challenging the conventional assumption of their equivalence.

0 citationsRead paper