Wildfire Detection Using Vision Transformer with the Wildfire Dataset

📅 2025-05-23
🏛️ 2025 Northeast Section Conference Proceedings
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
To address challenges in wildfire detection—including limited remote sensing coverage, smoke interference, difficulty in detecting small-scale fire targets, and poor model real-time performance—this paper proposes an end-to-end Vision Transformer (ViT)-based approach for wildfire image recognition. We present the first systematic adaptation of ViT architectures to wildfire detection, introducing novel data preprocessing and normalization strategies specifically designed for smoke-affected and low-contrast scenes. The model is trained on a high-resolution, 10.74 GB dataset comprising ‘fire’ and ‘nofire’ images. Evaluated on real-world wildfire data, it achieves 98.2% accuracy and 96.7% recall, with an F1-score 4.1 percentage points higher than ResNet-50. This demonstrates significantly improved robustness in early-stage fire detection and identification of small-scale flames. Moreover, the model supports efficient lightweight deployment on edge devices.

Technology Category

Application Category

📝 Abstract
The critical need for sophisticated detection techniques has been highlighted by the rising frequency and intensity of wildfires in the US, especially in California. In 2023, wildfires caused 130 deaths nationwide, the highest since 1990. In January 2025, Los Angeles wildfires which included the Palisades and Eaton fires burnt approximately 40,000 acres and 12,000 buildings, and caused loss of human lives. The devastation underscores the urgent need for effective detection and prevention strategies. Deep learning models, such as Vision Transformers (ViTs), can enhance early detection by processing complex image data with high accuracy. However, wildfire detection faces challenges, including the availability of high-quality, real-time data. Wildfires often occur in remote areas with limited sensor coverage, and environmental factors like smoke and cloud cover can hinder detection. Additionally, training deep learning models is computationally expensive, and issues like false positives/negatives and scaling remain concerns. Integrating detection systems with real-time alert mechanisms also poses difficulties. In this work, we used the wildfire dataset consisting of 10.74 GB high-resolution images categorized into 'fire' and 'nofire' classes is used for training the ViT model. To prepare the data, images are resized to 224 x 224 pixels, converted into tensor format, and normalized using ImageNet statistics.
Problem

Research questions and friction points this paper is trying to address.

Detecting wildfires early using Vision Transformers for accuracy
Addressing challenges in real-time wildfire data availability
Reducing false positives and scaling deep learning models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses Vision Transformer for wildfire detection
Processes high-resolution images for accuracy
Normalizes data with ImageNet statistics
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
G
Gowtham Raj Vuppari
Department of Computer Science and Engineering, University of Bridgeport
N
Navarun Gupta
Department of Electrical and Computer Engineering, University of Bridgeport
A
Ahmed El-Sayed
Department of Electrical and Computer Engineering, University of Bridgeport
X
Xingguo Xiong
Department of Electrical and Computer Engineering, University of Bridgeport