A Large-Scale Study of Model Integration in ML-Enabled Software Systems

📅 2024-08-12

🏛️ arXiv.org

📈 Citations: 0

✨ Influential: 0

career value

207K/year

🤖 AI Summary

This study addresses the challenges of integrating machine learning (ML) models into software systems—namely, poor integration practices, low reusability, and unclear architectural boundaries. It presents the first large-scale empirical investigation across 2,928 open-source ML-enabled systems. Leveraging GitHub code mining, static analysis, topic modeling, and architectural pattern identification, the work systematically characterizes ML integration topologies, code/model reuse practices, and maintenance bottlenecks. Key contributions include: (1) the first comprehensive classification framework and architectural pattern atlas for ML-enabled systems; (2) identification of seven prevalent integration topologies and four model reuse patterns; and (3) uncovering critical interdisciplinary collaboration barriers in ML-software co-development. The findings bridge the methodological gap between data science and software engineering at the model embedding stage, providing industry-practical architectural guidelines that significantly enhance the maintainability and reusability of ML systems.

Technology Category

Application Category

📝 Abstract

The rise of machine learning (ML) and its embedding in systems has drastically changed the engineering of software-intensive systems. Traditionally, software engineering focuses on manually created artifacts such as source code and the process of creating them, as well as best practices for integrating them, i.e., software architectures. In contrast, the development of ML artifacts, i.e. ML models, comes from data science and focuses on the ML models and their training data. However, to deliver value to end users, these ML models must be embedded in traditional software, often forming complex topologies. In fact, ML-enabled software can easily incorporate many different ML models. While the challenges and practices of building ML-enabled systems have been studied to some extent, beyond isolated examples, little is known about the characteristics of real-world ML-enabled systems. Properly embedding ML models in systems so that they can be easily maintained or reused is far from trivial. We need to improve our empirical understanding of such systems, which we address by presenting the first large-scale study of real ML-enabled software systems, covering over 2,928 open source systems on GitHub. We classified and analyzed them to determine their characteristics, as well as their practices for reusing ML models and related code, and the architecture of these systems. Our findings provide practitioners and researchers with insight into practices for embedding and integrating ML models, bringing data science and software engineering closer together.

Problem

Research questions and friction points this paper is trying to address.

ML model integration challenges

ML-enabled system characteristics

ML and code reuse practices

Innovation

Methods, ideas, or system contributions that make the work stand out.

Large-scale analysis of ML systems

Classification of ML integration patterns

Focus on model and code reuse

🔎 Similar Papers

No similar papers found.

💼 Related Jobs

Sr Machine Learning Engineer

Disney

The hiring range for this position in Burbank, CA is $155,700 - $208,700 per year and in Seattle is $163,100 - $218,700 per year. The base pay actually offered will take into account internal equity and also may vary depending on the candidate’s geographic region, job-related knowledge, skills, and experience among other factors. A bonus and/or long-term incentive units may be provided as part of the compensation package, in addition to the full range of medical, financial, and/or other benefits, dependent on the level and position offered.

Lake Buena Vista, FL, USA / USA - CA - 820 S Flower St, USA - FL - Kirkman Point 1, USA - WA - 925 4th Ave

Machine Learning Engineer