| Hauptseite > Publikationsdatenbank > TOAR-classifier v2: a data-driven classification tool for global air quality stations |
| Typ | Amount | VAT | Currency | Share | Status | Cost centre |
| APC | 1782.00 | 0.00 | EUR | 100.00 % | (Zahlung erfolgt) | ZB |
| Sum | 1782.00 | 0.00 | EUR | |||
| Total | 1782.00 |
| Journal Article | FZJ-2026-03445 |
; ; ; ;
2026
Copernicus
Katlenburg-Lindau
This record in other databases:
Please use a persistent id in citations: doi:10.5194/gmd-19-5765-2026 doi:10.34734/FZJ-2026-03445
Abstract: Accurate characterization of station locations is crucial for reliable air quality assessments such as the Tropospheric Ozone Assessment Report (TOAR). While urban and rural areas are relatively well-defined, the boundaries and identity of suburban areas remain ambiguous, overlapping with both urban and rural zones and varying due to cultural and social factors. This study investigates a machine learning approach to classify 24 348 stations in the unique global TOAR database as urban, suburban, or rural. We tested two different approaches: unsupervised K-means clustering with three clusters, and an ensemble of supervised learning classifiers including random forest, CatBoost, and LightGBM. We integrate these classifiers into a robust voting model, leveraging their collective predictive power. To address the inherent ambiguity of suburban areas, we implement a grid-search adjusted threshold probability technique. Our models, trained on the TOAR station metadata, are evaluated on 1979 unseen data points. K-means clustering achieves 71.88 % and 87.67 % accuracy for urban and rural areas respectively, but only 15.84 % for suburban zones. The supervised classifiers surpass this performance, reaching over 84 % accuracy for urban and rural categories, and 66 %–72 % for suburban areas. The adjusted threshold technique significantly enhances overall model accuracy, particularly for suburban classification. The good separation of our model is confirmed through evaluation with $NO_x$ and $PM_{2.5}$ concentration measurements, which were not included in the training data. Furthermore, manual inspection of 30 randomly selected sites with Google maps reveals that our method provides a better label for the station type than the labels that were reported by data providers and used in the model evaluation. The objective station classification proposed in this paper therefore provides a robust foundation for type-of-area-specific air quality assessments in TOAR and elsewhere.
|
The record appears in these collections: |