Institution profile

Florida Institute for Human and Machine Cognition

Academic institutionnorthamerica · us
Official website
Research library19linked papers
Opportunities0open roles
Selected work

Representative Papers

SeePerSea: Multi-modal Perception Dataset of In-water Objects for Autonomous Surface Vehicles

Apr 29, 2024

A lack of publicly available, high-quality underwater obstacle perception datasets for autonomous surface vehicles (ASVs) operating in complex hydrodynamic environments hinders progress in marine autonomous navigation. Method: This work introduces ASV-Underwater—the first open-source, multimodal underwater obstacle dataset specifically designed for ASVs. Collected over four years, it encompasses diverse targets, turbid water conditions, low-light scenarios, and dynamic occlusions, with temporally synchronized optical and acoustic imagery. All data are annotated with ego-centric, fine-grained pixel-level and bounding-box labels, and formatted to comply with standard detection frameworks (e.g., YOLOv8, Faster R-CNN). Contribution/Results: Experiments demonstrate that models trained on ASV-Underwater achieve significantly improved robustness in underwater obstacle detection and classification, effectively addressing a critical data gap in maritime perception research.

1 citationsRead paper

HANDOFF: Humanoid Agentic Task-Space Whole-Body Control via Distilled Complementary Teachers

Jun 04, 2026

This work addresses the challenge of semantic mismatch between high-level task instructions and low-level whole-body control commands in real-world humanoid robot deployment, where existing approaches struggle to generate robust motions from natural language. The authors propose HANDOFF, a unified whole-body controller built upon a compact, explicit task-space interface that enables end-to-end mapping from natural language instructions to robust full-body actions without task-specific fine-tuning. HANDOFF introduces context-gated multi-teacher KL distillation to seamlessly integrate three expert policies—motion tracking, walking, and fall recovery—within a conditional gating mechanism and mixture-of-experts architecture, augmented with safety-aware filtering during training. Evaluated on the Unitree G1 platform, the method achieves state-of-the-art linear velocity tracking accuracy and an expansive robust operating range, successfully demonstrating diverse instruction-following tasks on real hardware.

0 citationsRead paper

Generating Satellite Imagery Data for Wildfire Detection through Mask-Conditioned Generative AI

Apr 02, 2026

This study addresses the scarcity of labeled satellite imagery for wildfire monitoring by introducing, for the first time, the mask-conditional diffusion model EarthSynth to synthesize post-disaster Sentinel-2 RGB images. Leveraging an image inpainting architecture combined with structured prompting strategies—including handwritten cues and vision-language prompts generated by Qwen2-VL—the method enables controllable generation without task-specific fine-tuning. A region-aware color-matching post-processing step further enhances photorealism. Experimental results demonstrate that the inpainting paradigm significantly outperforms full-image generation across four key metrics: Burn IoU (peaking at 0.456), color distance, darkness contrast, and spectral plausibility, offering an effective data augmentation solution for wildfire detection.

0 citationsRead paper
Recent publications

Latest Papers

HANDOFF: Humanoid Agentic Task-Space Whole-Body Control via Distilled Complementary Teachers

Jun 04, 2026

This work addresses the challenge of semantic mismatch between high-level task instructions and low-level whole-body control commands in real-world humanoid robot deployment, where existing approaches struggle to generate robust motions from natural language. The authors propose HANDOFF, a unified whole-body controller built upon a compact, explicit task-space interface that enables end-to-end mapping from natural language instructions to robust full-body actions without task-specific fine-tuning. HANDOFF introduces context-gated multi-teacher KL distillation to seamlessly integrate three expert policies—motion tracking, walking, and fall recovery—within a conditional gating mechanism and mixture-of-experts architecture, augmented with safety-aware filtering during training. Evaluated on the Unitree G1 platform, the method achieves state-of-the-art linear velocity tracking accuracy and an expansive robust operating range, successfully demonstrating diverse instruction-following tasks on real hardware.

0 citationsRead paper

Generating Satellite Imagery Data for Wildfire Detection through Mask-Conditioned Generative AI

Apr 02, 2026

This study addresses the scarcity of labeled satellite imagery for wildfire monitoring by introducing, for the first time, the mask-conditional diffusion model EarthSynth to synthesize post-disaster Sentinel-2 RGB images. Leveraging an image inpainting architecture combined with structured prompting strategies—including handwritten cues and vision-language prompts generated by Qwen2-VL—the method enables controllable generation without task-specific fine-tuning. A region-aware color-matching post-processing step further enhances photorealism. Experimental results demonstrate that the inpainting paradigm significantly outperforms full-image generation across four key metrics: Burn IoU (peaking at 0.456), color distance, darkness contrast, and spectral plausibility, offering an effective data augmentation solution for wildfire detection.

0 citationsRead paper

Embedding Classical Balance Control Principles in Reinforcement Learning for Humanoid Recovery

Mar 09, 2026

This work addresses the challenge of fall recovery for humanoid robots in unstructured environments, where existing reinforcement learning approaches often lack explicit modeling of balance dynamics. The authors propose a novel method that incorporates classical balance metrics—capture point, center-of-mass state, and centroidal momentum—as privileged inputs to the critic and leverages them to design shaping rewards. Relying solely on proprioceptive feedback, the approach achieves zero-shot transfer to real hardware. By embedding interpretable principles of balance control, the method learns a single, physically consistent recovery policy that generalizes across the full spectrum of disturbances—from minor perturbations to multi-contact falls. Evaluated on the Unitree H1-2 platform, the policy attains a 93.4% success rate in random fall recovery. Ablation studies confirm the critical role of the balance-aware architecture in policy learning, with successful demonstrations in both Sim-to-Sim transfer and preliminary real-world deployment.

0 citationsRead paper