Institution profile

dida Datenschmiede GmbH

Industry researcheurope · de
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Reinforcement Learning with Random Time Horizons

Jun 01, 2025

This paper addresses realistic reinforcement learning scenarios where trajectory termination times are both stochastic and policy-dependent—challenging the conventional assumptions of fixed or infinite horizons. Method: We develop a dual modeling framework integrating trajectory- and state-space perspectives, grounded in stochastic process theory and optimal control principles, and propose a Monte Carlo–based gradient estimation scheme. Contribution/Results: (1) We derive the first unbiased policy gradient theorem for policy-dependent stochastic horizons; (2) we unify the theoretical foundations of policy optimization across finite-, infinite-, and stochastic-horizon settings; and (3) empirical evaluations demonstrate that the proposed gradient estimator significantly improves convergence speed and training stability compared to classical methods—including those assuming fixed or discounted infinite horizons—thereby enabling more robust and efficient learning in real-world sequential decision-making problems with uncertain termination.

0 citationsRead paper

Meta-learning For Few-Shot Time Series Crop Type Classification: A Benchmark On The EuroCropsML Dataset

Apr 15, 2025

To address geographic data imbalance and poor generalization under few-shot conditions in remote sensing–based crop classification, this paper introduces the first meta-learning benchmark tailored to real-world agricultural scenarios, built upon the multi-national EuroCropsML time-series satellite dataset. We systematically evaluate few-shot meta-learning methods—including FO-MAML, ANIL, and TIML—as well as cross-domain transfer learning for regional generalization, integrating Sentinel-2 time-series modeling with farmer-reported ground-truth validation. Key findings: (1) Geographic distance imposes a significant constraint on knowledge transfer; (2) A pronounced accuracy–computational cost trade-off is quantified—e.g., MAML-based methods yield marginal accuracy gains (+0.5%) on Estonian tasks but incur substantial computational overhead; (3) Cross-domain transfer (e.g., Estonia ↔ Portugal) suffers severe performance degradation; (4) We publicly release the benchmark codebase and standardized evaluation protocol.

0 citationsRead paper

Underdamped Diffusion Bridges with Applications to Sampling

Mar 02, 2025

This work addresses the challenge of efficient sampling from unnormalized densities without access to target samples. Methodologically, it introduces the Underdamped Diffusion Bridge (UDB) framework: (i) it establishes, for the first time under underdamped stochastic differential equations (SDEs), the rigorous equivalence between score matching and maximizing a variational lower bound; (ii) it abandons fixed noise schedules and instead learns a flexible, task-agnostic density evolution path; and (iii) it incorporates a degenerate diffusion matrix and high-order numerical integrators to enhance stability and accuracy. The key contributions are: (i) an end-to-end sampling scheme from a prior to an unnormalized target distribution—requiring neither target samples nor hyperparameter tuning; and (ii) state-of-the-art performance across diverse no-sample tasks, with significantly reduced sampling steps, while maintaining theoretical rigor and practical deployability.

0 citationsRead paper

From discrete-time policies to continuous-time diffusion samplers: Asymptotic equivalences and faster training

Jan 10, 2025

This work addresses efficient training of diffusion models under target-free sample conditions. Conventional approaches either rely on target examples or incur prohibitive computational costs. To overcome these limitations, we propose an unsupervised training framework grounded in neural stochastic differential equations (SDEs). We first establish a rigorous theoretical equivalence—under infinitesimal time steps—between the entropy-regularized reinforcement learning objective of GFlowNets and both the continuous-time Fokker–Planck equation and the path-space variational objective. Leveraging this insight, we design a coarse-grained temporal discretization scheme that enables time-local optimization while preserving asymptotic consistency. Empirically, our method achieves state-of-the-art performance on standard sampling benchmarks, significantly improving sample efficiency and reducing computational cost by approximately 40%–60%. This work introduces a theoretically grounded and practically efficient paradigm for unsupervised generative modeling.

0 citationsRead paper
Recent publications

Latest Papers

Reinforcement Learning with Random Time Horizons

Jun 01, 2025

This paper addresses realistic reinforcement learning scenarios where trajectory termination times are both stochastic and policy-dependent—challenging the conventional assumptions of fixed or infinite horizons. Method: We develop a dual modeling framework integrating trajectory- and state-space perspectives, grounded in stochastic process theory and optimal control principles, and propose a Monte Carlo–based gradient estimation scheme. Contribution/Results: (1) We derive the first unbiased policy gradient theorem for policy-dependent stochastic horizons; (2) we unify the theoretical foundations of policy optimization across finite-, infinite-, and stochastic-horizon settings; and (3) empirical evaluations demonstrate that the proposed gradient estimator significantly improves convergence speed and training stability compared to classical methods—including those assuming fixed or discounted infinite horizons—thereby enabling more robust and efficient learning in real-world sequential decision-making problems with uncertain termination.

0 citationsRead paper

Meta-learning For Few-Shot Time Series Crop Type Classification: A Benchmark On The EuroCropsML Dataset

Apr 15, 2025

To address geographic data imbalance and poor generalization under few-shot conditions in remote sensing–based crop classification, this paper introduces the first meta-learning benchmark tailored to real-world agricultural scenarios, built upon the multi-national EuroCropsML time-series satellite dataset. We systematically evaluate few-shot meta-learning methods—including FO-MAML, ANIL, and TIML—as well as cross-domain transfer learning for regional generalization, integrating Sentinel-2 time-series modeling with farmer-reported ground-truth validation. Key findings: (1) Geographic distance imposes a significant constraint on knowledge transfer; (2) A pronounced accuracy–computational cost trade-off is quantified—e.g., MAML-based methods yield marginal accuracy gains (+0.5%) on Estonian tasks but incur substantial computational overhead; (3) Cross-domain transfer (e.g., Estonia ↔ Portugal) suffers severe performance degradation; (4) We publicly release the benchmark codebase and standardized evaluation protocol.

0 citationsRead paper

Underdamped Diffusion Bridges with Applications to Sampling

Mar 02, 2025

This work addresses the challenge of efficient sampling from unnormalized densities without access to target samples. Methodologically, it introduces the Underdamped Diffusion Bridge (UDB) framework: (i) it establishes, for the first time under underdamped stochastic differential equations (SDEs), the rigorous equivalence between score matching and maximizing a variational lower bound; (ii) it abandons fixed noise schedules and instead learns a flexible, task-agnostic density evolution path; and (iii) it incorporates a degenerate diffusion matrix and high-order numerical integrators to enhance stability and accuracy. The key contributions are: (i) an end-to-end sampling scheme from a prior to an unnormalized target distribution—requiring neither target samples nor hyperparameter tuning; and (ii) state-of-the-art performance across diverse no-sample tasks, with significantly reduced sampling steps, while maintaining theoretical rigor and practical deployability.

0 citationsRead paper

From discrete-time policies to continuous-time diffusion samplers: Asymptotic equivalences and faster training

Jan 10, 2025

This work addresses efficient training of diffusion models under target-free sample conditions. Conventional approaches either rely on target examples or incur prohibitive computational costs. To overcome these limitations, we propose an unsupervised training framework grounded in neural stochastic differential equations (SDEs). We first establish a rigorous theoretical equivalence—under infinitesimal time steps—between the entropy-regularized reinforcement learning objective of GFlowNets and both the continuous-time Fokker–Planck equation and the path-space variational objective. Leveraging this insight, we design a coarse-grained temporal discretization scheme that enables time-local optimization while preserving asymptotic consistency. Empirically, our method achieves state-of-the-art performance on standard sampling benchmarks, significantly improving sample efficiency and reducing computational cost by approximately 40%–60%. This work introduces a theoretically grounded and practically efficient paradigm for unsupervised generative modeling.

0 citationsRead paper