Institution profile

Dotphoton

Industry researcheurope · ch
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Making the Discrete Continuous: Synthetic RAW Augmentations for Fine-Grained Evaluation of Person Detection Performance in Low Light

May 21, 2026

This work addresses the scarcity and uneven distribution of real-world pedestrian detection data under low-light conditions, which hinders fine-grained performance evaluation. To bridge this gap, the study introduces— for the first time—a RAW image synthesis method grounded in the physical noise model of camera sensors, enabling continuous expansion of the low-light input space to generate high-fidelity synthetic samples. This approach substantially enhances dataset coverage for benchmarking and demonstrates strong performance alignment between synthetic and real low-light data across multiple state-of-the-art object detectors. By effectively closing the evaluation gap, the proposed framework establishes a reliable and scalable paradigm for assessing pedestrian detection in low-light scenarios.

0 citationsRead paper

Croissant Baker: Metadata Generation for Discoverable, Governable, and Reusable ML Datasets

May 14, 2026

This work addresses the limitations of existing Croissant metadata generation approaches, which rely on public platforms and struggle to accommodate governed or large-scale local datasets. The authors propose the first open-source, local-first command-line tool that directly generates Croissant-compliant JSON-LD metadata from local directories via a modular processor registration mechanism, supporting mainstream formats such as Parquet. By eliminating dependence on external platforms, this method significantly enhances the discoverability and reusability of private, high-value datasets. Experimental evaluation across more than 140 datasets—including MIMIC-IV with 886 million rows—demonstrates that the generated metadata achieves 97–100% accuracy, matching or exceeding that of manual curation or standard methods.

0 citationsRead paper

Autoguided Online Data Curation for Diffusion Model Training

Sep 18, 2025

This study addresses the time and sample efficiency bottlenecks in training generative diffusion models. We propose a unified framework integrating autoguidance and Joint Example Selection for Training (JEST), enabling online data filtering and dynamic optimization during training. Through controlled experiments on 2D synthetic data and 64×64 image generation, we systematically evaluate how different data selection strategies affect generation quality and diversity. Results show that autoguidance consistently improves both sample fidelity and diversity; early-stage AJEST achieves comparable or slightly better data efficiency than autoguidance but suffers from high computational overhead, limiting practical deployment. Our key contribution is the first empirical delineation of the stage-dependent effectiveness boundaries between autoguidance and JEST—revealing that lightweight autoguidance dominates across most training phases in terms of both performance and deployment feasibility.

0 citationsRead paper

AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Feb 19, 2025arXiv.org

The absence of standardized benchmarks for evaluating the safety and robustness of AI systems under adversarial prompts hinders rigorous, comparable assessments. Method: This paper introduces AILuminate v1.0—the first industry-grade benchmark for AI risk and reliability—covering 12 harm categories: brute-force attacks, criminal activity, child exploitation, suicide/self-harm, intellectual property infringement, privacy violations, defamation, hate speech, sexually explicit content, and domain-specific risks (elections, finance, health, law). It proposes a novel five-level interpretable scoring framework and an entropy-based quantification method for response quality, alongside a large-scale adversarial prompt dataset enabling single-turn evaluation. Contribution/Results: As the first fully multidimensional, reproducible, and extensible AI safety benchmark, AILuminate v1.0 supports open, collaborative evolution and provides empirically grounded tools for developers, deployers, and policymakers to advance global AI safety standardization.

0 citationsRead paper
Recent publications

Latest Papers

Making the Discrete Continuous: Synthetic RAW Augmentations for Fine-Grained Evaluation of Person Detection Performance in Low Light

May 21, 2026

This work addresses the scarcity and uneven distribution of real-world pedestrian detection data under low-light conditions, which hinders fine-grained performance evaluation. To bridge this gap, the study introduces— for the first time—a RAW image synthesis method grounded in the physical noise model of camera sensors, enabling continuous expansion of the low-light input space to generate high-fidelity synthetic samples. This approach substantially enhances dataset coverage for benchmarking and demonstrates strong performance alignment between synthetic and real low-light data across multiple state-of-the-art object detectors. By effectively closing the evaluation gap, the proposed framework establishes a reliable and scalable paradigm for assessing pedestrian detection in low-light scenarios.

0 citationsRead paper

Croissant Baker: Metadata Generation for Discoverable, Governable, and Reusable ML Datasets

May 14, 2026

This work addresses the limitations of existing Croissant metadata generation approaches, which rely on public platforms and struggle to accommodate governed or large-scale local datasets. The authors propose the first open-source, local-first command-line tool that directly generates Croissant-compliant JSON-LD metadata from local directories via a modular processor registration mechanism, supporting mainstream formats such as Parquet. By eliminating dependence on external platforms, this method significantly enhances the discoverability and reusability of private, high-value datasets. Experimental evaluation across more than 140 datasets—including MIMIC-IV with 886 million rows—demonstrates that the generated metadata achieves 97–100% accuracy, matching or exceeding that of manual curation or standard methods.

0 citationsRead paper

Autoguided Online Data Curation for Diffusion Model Training

Sep 18, 2025

This study addresses the time and sample efficiency bottlenecks in training generative diffusion models. We propose a unified framework integrating autoguidance and Joint Example Selection for Training (JEST), enabling online data filtering and dynamic optimization during training. Through controlled experiments on 2D synthetic data and 64×64 image generation, we systematically evaluate how different data selection strategies affect generation quality and diversity. Results show that autoguidance consistently improves both sample fidelity and diversity; early-stage AJEST achieves comparable or slightly better data efficiency than autoguidance but suffers from high computational overhead, limiting practical deployment. Our key contribution is the first empirical delineation of the stage-dependent effectiveness boundaries between autoguidance and JEST—revealing that lightweight autoguidance dominates across most training phases in terms of both performance and deployment feasibility.

0 citationsRead paper

AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons

Feb 19, 2025arXiv.org

The absence of standardized benchmarks for evaluating the safety and robustness of AI systems under adversarial prompts hinders rigorous, comparable assessments. Method: This paper introduces AILuminate v1.0—the first industry-grade benchmark for AI risk and reliability—covering 12 harm categories: brute-force attacks, criminal activity, child exploitation, suicide/self-harm, intellectual property infringement, privacy violations, defamation, hate speech, sexually explicit content, and domain-specific risks (elections, finance, health, law). It proposes a novel five-level interpretable scoring framework and an entropy-based quantification method for response quality, alongside a large-scale adversarial prompt dataset enabling single-turn evaluation. Contribution/Results: As the first fully multidimensional, reproducible, and extensible AI safety benchmark, AILuminate v1.0 supports open, collaborative evolution and provides empirically grounded tools for developers, deployers, and policymakers to advance global AI safety standardization.

0 citationsRead paper