Safety-Oriented Air Quality Index Classification for Imbalanced Data Using Optimized Boosting Models with Optuna and Oversampling
DOI:
https://doi.org/10.38043/tiers.v7i1.7558Keywords:
Air Quality Index, Extreme Class Imbalance, Ensemble Boosting, SMOTE, Minority Class Detection, Robustness EvaluationAbstract
Air Quality Index (AQI) classification is essential for communicating environmental health risks. However, hazardous air conditions occur far less frequently than normal conditions, challenging conventional classification models. This study investigates multi-class AQI classification using the "Global Air Quality 2025" dataset, comprising 52,704 observations with an extreme class imbalance ratio of approximately 1:173. Under such conditions, conventional accuracy metrics often mask systemic failures in detecting critical minority classes. To address potential data leakage present in previous approaches, this research implements a rigorous cross-validation architecture combined with an independent 20% hold-out test set. The methodology employs an Ablation Study to systematically isolate the impacts of Optuna hyperparameter tuning guided by Macro F1-Score and oversampling techniques (SMOTE and ADASYN). The results demonstrate that the proposed Hybrid-SMOTE LightGBM configuration successfully balances hazard detection sensitivity with global stability. On the unseen hold-out set, the optimal model achieved a Macro F1-Score of 0.8079, an accuracy of 92.80%, and a ROC-AUC of 0.9847. Crucially, the model delivered a 65.12% recall for the critical Unhealthy minority class, a nearly 40% improvement over the baseline. Error profile analysis confirmed the model's safety-oriented robustness, as 97.6% of peak hazardous events were either accurately classified or safely constrained to the adjacent warning category, minimizing catastrophic misclassifications. These findings prove that reliable detection of environmental hazards requires safety-oriented per-class evaluation and strict validation frameworks, as reliance on aggregate global metrics leads to dangerously misleading performance assessments.
Downloads
References
O. Sokolova, A. Yurgenson, and V. Shakhov, Development of Air Quality Monitoring Systems: Balancing Infrastructure Investment and User Satisfaction Policies, Sensors, vol. 25, no. 3, p. 875, Jan. 2025, doi: 10.3390/s25030875.
Z. Zhong et al., Dynamic associations between long-term exposure to ambient air pollution and respiratory-cardiovascular diseases: A trajectory analysis of a prospective study, Ecotoxicol. Environ. Saf., vol. 306, p. 119329, Nov. 2025, doi: 10.1016/j.ecoenv.2025.119329.
M. M. Abdelmalek, H. Mahmoud, and H. Shokry, Prognosis of air quality index and air pollution using machine learning techniques, Sci. Rep., vol. 15, no. 1, p. 25890, Jul. 2025, doi: 10.1038/s41598-025-11260-y.
V. R. Folifack Signing et al., IoT-based monitoring system and air quality prediction using machine learning for a healthy environment in Cameroon, Environ. Monit. Assess., vol. 196, no. 7, p. 621, Jul. 2024, doi: 10.1007/s10661-024-12789-7.
R. S. Rao, L. R. Kalabarige, B. Alankar, and A. K. Sahu, Multimodal imputation-based stacked ensemble for prediction and classification of air quality index in Indian cities, Computers and Electrical Engineering, vol. 114, p. 109098, Mar. 2024, doi: 10.1016/j.compeleceng.2024.109098.
A. J. Barid, H. Hadiyanto, and A. Wibowo, Optimization of the algorithms use ensemble and synthetic minority oversampling technique for air quality classification, Indonesian Journal of Electrical Engineering and Computer Science, vol. 33, no. 3, p. 1632, Mar. 2024, doi: 10.11591/ijeecs.v33.i3.pp1632-1640.
T. Toharudin et al., Boosting Algorithm to Handle Unbalanced Classification of PM2.5Concentration Levels by Observing Meteorological Parameters in Jakarta-Indonesia Using AdaBoost, XGBoost, CatBoost, and LightGBM, IEEE Access, vol. 11, pp. 3568035696, 2023, doi: 10.1109/ACCESS.2023.3265019.
I. Tanasa, M. Cazacu, and B. Sluser, Air Quality Integrated Assessment: Environmental Impacts, Risks and Human Health Hazards, Applied Sciences, vol. 13, no. 2, p. 1222, Jan. 2023, doi: 10.3390/app13021222.
N. Masseran, M. A. M. Safari, and R. R. M. Tajuddin, Probabilistic classification of the severity classes of unhealthy air pollution events, Environ. Monit. Assess., vol. 196, no. 6, p. 523, Jun. 2024, doi: 10.1007/s10661-024-12700-4.
G. Velarde et al., Tree boosting methods for balanced and imbalanced classification and their robustness over time in risk assessment, Intelligent Systems with Applications, vol. 22, p. 200354, Jun. 2024, doi: 10.1016/j.iswa.2024.200354.
S. Ketu and P. K. Mishra, Scalable kernel-based SVM classification algorithm on imbalance air quality data for proficient healthcare, Complex & Intelligent Systems, vol. 7, no. 5, pp. 25972615, Oct. 2021, doi: 10.1007/s40747-021-00435-5.
Z.-Y. Chen, H. Petetin, R. F. Mndez Turrubiates, H. Achebak, C. Prez Garca-Pando, and J. Ballester, Population exposure to multiple air pollutants and its compound episodes in Europe, Nat. Commun., vol. 15, no. 1, p. 2094, Mar. 2024, doi: 10.1038/s41467-024-46103-3.
W. Chandra, B. Suprihatin, and Y. Resti, Median-KNN Regressor-SMOTE-Tomek Links for Handling Missing and Imbalanced Data in Air Quality Prediction, Symmetry (Basel)., vol. 15, no. 4, p. 887, Apr. 2023, doi: 10.3390/sym15040887.
H. Alkabbani, A. Ramadan, Q. Zhu, and A. Elkamel, An Improved Air Quality Index Machine Learning-Based Forecasting with Multivariate Data Imputation Approach, Atmosphere (Basel)., vol. 13, no. 7, p. 1144, Jul. 2022, doi: 10.3390/atmos13071144.
M. Karmoude et al., Machine learning for air quality prediction and data analysis: Review on recent advancements, challenges, and outlooks, Science of The Total Environment, vol. 1002, p. 180593, Nov. 2025, doi: 10.1016/j.scitotenv.2025.180593.
S. P. Praveen et al., Enhanced feature selection and ensemble learning for cardiovascular disease prediction: hybrid GOL2-2T and adaptive boosted decision fusion with babysitting refinement, Front. Med. (Lausanne)., vol. 11, Jul. 2024, doi: 10.3389/fmed.2024.1407376.
K. Abnoosian, R. Farnoosh, and M. H. Behzadi, Prediction of diabetes disease using an ensemble of machine learning multi-classifier models, BMC Bioinformatics, vol. 24, no. 1, p. 337, Sep. 2023, doi: 10.1186/s12859-023-05465-z.
M. Mujahid et al., Data oversampling and imbalanced datasets: an investigation of performance for machine learning and feature engineering, J. Big Data, vol. 11, no. 1, p. 87, Jun. 2024, doi: 10.1186/s40537-024-00943-4.
Haibo He, Yang Bai, E. A. Garcia, and Shutao Li, ADASYN: Adaptive synthetic sampling approach for imbalanced learning, in 2008 IEEE International Joint Conference on Neural Networks (IEEE World Congress on Computational Intelligence), IEEE, Jun. 2008, pp. 13221328. doi: 10.1109/IJCNN.2008.4633969.
Y.-W. Wang, H.-C. Yang, Y.-H. Chen, and C.-Y. Guo, Generalized Estimating Equations Boosting (GEEB) machine for correlated data, J. Big Data, vol. 11, no. 1, p. 20, Jan. 2024, doi: 10.1186/s40537-023-00875-5.
S. Saminathan and C. Malathy, Ensemble-based classification approach for PM2.5 concentration forecasting using meteorological data, Front. Big Data, vol. 6, Jun. 2023, doi: 10.3389/fdata.2023.1175259.
X. Zhang, X. Jiang, and Y. Li, Prediction of air quality index based on the SSA-BiLSTM-LightGBM model, Sci. Rep., vol. 13, no. 1, p. 5550, Apr. 2023, doi: 10.1038/s41598-023-32775-2.
J. Ugirumurera, E. A. Bensen, J. Severino, and J. Sanyal, Addressing bias in bagging and boosting regression models, Sci. Rep., vol. 14, no. 1, p. 18452, Aug. 2024, doi: 10.1038/s41598-024-68907-5.
M. S. Sawah, H. Elmannai, A. A. El-Bary, Kh. Lotfy, and O. E. Sheta, Improving air quality prediction using hybrid BPSO with BWAO for feature selection and hyperparameters optimization, Sci. Rep., vol. 15, no. 1, p. 13176, Apr. 2025, doi: 10.1038/s41598-025-95983-y.
H. Shao, X. Liu, D. Zong, and Q. Song, Optimization of diabetes prediction methods based on combinatorial balancing algorithm, Nutr. Diabetes, vol. 14, no. 1, p. 63, Aug. 2024, doi: 10.1038/s41387-024-00324-z.
D. Elreedy, A. F. Atiya, and F. Kamalov, A theoretical distribution analysis of synthetic minority oversampling technique (SMOTE) for imbalanced learning, Mach. Learn., vol. 113, no. 7, pp. 49034923, Jul. 2024, doi: 10.1007/s10994-022-06296-4.
S. Farhadpour, T. A. Warner, and A. E. Maxwell, Selecting and Interpreting Multiclass Loss and Accuracy Assessment Metrics for Classifications with Class Imbalance: Guidance and Best Practices, Remote Sens. (Basel)., vol. 16, no. 3, p. 533, Jan. 2024, doi: 10.3390/rs16030533.
S. Sheikholeslami, M. Meister, T. Wang, A. H. Payberah, V. Vlassov, and J. Dowling, AutoAblation: Automated Parallel Ablation Studies for Deep Learning, in Proceedings of the 1st Workshop on Machine Learning and Systems, New York, NY, USA: ACM, Apr. 2021, pp. 5561. doi: 10.1145/3437984.3458834.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Made Yudi Dwipayana, Gede Angga Pradipta, Dandy Pramana Hostiadi

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.