International Journal of Technology and Emerging Research

DOI: 10.64823/ijter.2504006

⚠️ This HTML version is automatically generated from the manuscript file and may contain formatting or data discrepancies compared to the original paper. Please refer to the PDF version for the authoritative, publisher-formatted record.

Comparative Statistical Inference of PM2.5 Levels Across Indian Cities : A Bootstrap vs Classical Approach

Dr. Y. Raghunatha Reddy, Coordinator, Dept. of OR&SQC, Rayalaseema University Kurnool. drraghuy@gmail.com

B. Sravanthi, Asst. Prof, Dept. of H&S, G.Pullaiah College of Engg.& Tech, Kurnool

S. Rehana, Lecturer, Dept. of Statistics, Ravindra Degree College for Women, Kurnool

Abstract
Air pollution remains a pressing environmental and public health challenge in India, with fine particulate matter (PM2.5) posing severe respiratory and cardiovascular risks. This study conducts a comparative statistical inference analysis of daily PM2.5 concentrations for Delhi and Mumbai, based on 2024 data sourced from the Central Pollution Control Board (CPCB). Two estimation approaches are applied: the classical parametric t-based confidence interval method, which assumes normality, and the non-parametric bootstrap approach, which relies on re-sampling without distributional assumptions. The analysis reveals that while Delhi consistently exhibits substantially higher PM2.5 levels than Mumbai, the estimated means and confidence intervals from both methods are closely aligned, indicating that the parametric method’s assumptions are reasonably met in this dataset. The findings underscore the utility of bootstrap methods in validating classical inference, particularly in environmental data analysis, and provide robust evidence for policy-oriented air quality interventions.

1. Introduction

Air pollution is one of the most critical environmental and public health challenges in contemporary India, with fine particulate matter (PM2.5) being recognized as a particularly harmful pollutant. PM2.5 refers to airborne particles with a diameter less than or equal to 2.5 micrometers, small enough to penetrate deep into the alveolar regions of the lungs and even enter the bloodstream. Chronic exposure to elevated PM2.5 levels has been linked to respiratory diseases, cardiovascular disorders, reduced life expectancy, and increased mortality rates.

Urban centers such as Delhi and Mumbai represent contrasting yet significant case studies for understanding the scale and variability of PM2.5 pollution in India. Delhi, located in the Indo-Gangetic plain, is frequently ranked among the most polluted cities in the world due to a combination of vehicular emissions, industrial activity, biomass burning, and unfavorable meteorological conditions. In contrast, Mumbai, a coastal metropolis, benefits from sea breezes and higher humidity, which can disperse pollutants more effectively — though the city still faces periodic spikes in pollution due to industrial zones, construction activity, and seasonal weather patterns.

Accurate estimation of PM2.5 levels, along with quantification of their uncertainty, is vital for designing effective environmental policies and intervention strategies. Statistical inference provides a formal framework for such estimation, enabling researchers to make generalizable conclusions about the population from a sample of observed data.

2. Dataset and Methods

2.1 Dataset Description

The dataset used in this study contains daily average PM2.5 concentrations (in micrograms per cubic meter, μg/m³) for Delhi and Mumbai covering the period January 1, 2024 to June 30, 2024. The data originates from the Central Pollution Control Board (CPCB), the apex governmental body in India responsible for monitoring and regulating air quality.

Daily PM2.5 values were computed by aggregating hourly readings from multiple monitoring stations within each city. To maintain data integrity:

The dataset size consists of 182 observations per city, ensuring a reasonably large sample for statistical inference.

2.2 Statistical Methods:

The study compares two approaches for constructing confidence intervals (CIs) for the mean PM2.5 concentration:

2.2.1 Classical t-based Confidence Interval

2.2.2 Bootstrap Confidence Interval

2.3 Comparative Framework

By applying both methods to the same dataset, the analysis assesses:

3. Assumptions and Limitations

3.1 Assumptions

  1. Representativeness and accuracy of Data
  2. Independence of Observations
  3. Normality for Classical Inference
  4. Random Sampling for Bootstrap

3.2 Limitations

  1. Temporal Scope
  2. Geographical Coverage
  3. Potential Measurement Bias
  4. Limited Variable Scope
  5. Statistical Limitations

4. Statistical analysis:

4.1 Descriptive Statistics

The following table presents the summary statistics of daily average PM2.5 concentrations for Delhi and Mumbai from January 1, 2024 to June 30, 2024.

City

N

Mean (μg/m³)

Std. Dev.

Minimum

Maximum

Median

IQR

Delhi

182

112.4

28.7

62.3

192.6

109.8

34.5

Mumbai

182

54.8

16.1

27.4

92.1

53.1

19.8

Delhi’s PM2.5 levels are more than twice those of Mumbai on average, with a higher variability, indicating more frequent extreme pollution days. Mumbai’s distribution is narrower, reflecting greater stability in air quality conditions.

4.2 Classical t-based Confidence Intervals

For each city, a 95% confidence interval (CI) for the mean PM2.5 concentration was calculated using the t-distribution.

Even accounting for sampling variability, there is no overlap between the CIs for Delhi and Mumbai, indicating a statistically significant difference in mean PM2.5 levels.

4.3 Bootstrap Confidence Intervals

Using 5,000 bootstrap resamples, the percentile method was applied:

The bootstrap intervals closely match the t-based intervals, suggesting that the assumption of normality for the mean is reasonable for this dataset.

4.4 Hypothesis Testing

To formally test the difference between the two cities’ PM2.5 levels:

H0: there is no significant difference in average PM2.5 levels between Delhi and Mumbai

Using a two-sample t-test (assuming unequal variances):

The difference in average PM2.5 levels between Delhi and Mumbai is highly statistically significant.

4.5 Exploratory Data Analysis (EDA)

The exploratory data analysis aims to understand the structure, patterns, and variability of the 2024 PM2.5 data from Delhi and Mumbai before applying formal inference.

4.5.1 Time Series Overview

C:\Users\User1PC\Downloads\figure1_timeseries.png

Figure 1: Daily PM2.5 concentrations for Delhi and Mumbai.

4.5.2 Distribution Analysis

C:\Users\User1PC\Downloads\figure3_histogram.pngHistograms and Kernel Density Estimates (KDEs) reveal:

Figure 2: Histograms for both cities indicate approximately symmetric distributions with slight right skew, more pronounced in Delhi.

4.5.3 Boxplot Comparison

C:\Users\User1PC\Downloads\figure2_boxplot.pngBoxplots highlights:

Figure 3:: Boxplots of PM2.5 concentrations show clear elevation and wider spread for Delhi

4.5.4 Correlation with Time

5. Conclusions

This study applied both classical parametric inference and non-parametric bootstrap methods to assess PM2.5 air pollution levels in Delhi and Mumbai during the first half of 2024, using official data from the Central Pollution Control Board (CPCB).

The exploratory data analysis (EDA) revealed stark differences in air quality between the two cities:

The inferential results were consistent across both estimation frameworks:

From a methodological standpoint:

  1. Bootstrap methods proved valuable for validating classical inference results, especially in the presence of skewness and outliers, as observed in Delhi’s data.
  2. The similarity of the two approaches in this case suggests robustness of the findings.

From an environmental policy perspective:

This analysis not only provides empirical evidence of the alarming state of urban air pollution in India’s major cities but also demonstrates the complementary use of classical and bootstrap inference techniques in environmental statistics. The approach can be replicated for other pollutants, cities, or time periods to support data-driven policymaking.

6. References

  1. Central Pollution Control Board (CPCB). (2024). National Air Quality Monitoring Programme (NAMP) – PM2.5 Data. Ministry of Environment, Forest and Climate Change, Government of India. Retrieved from: https://cpcb.nic.in
  2. World Health Organization (WHO). (2021). WHO global air quality guidelines: Particulate matter (PM₂.₅ and PM₁₀), ozone, nitrogen dioxide, sulfur dioxide and carbon monoxide. Geneva: WHO. Retrieved from: https://www.who.int/publications/i/item/9789240034228
  3. Davison, A. C., & Hinkley, D. V. (1997). Bootstrap Methods and Their Application. Cambridge: Cambridge University Press. doi:10.1017/CBO9780511802843
  4. Efron, B., & Tibshirani, R. J. (1994). An Introduction to the Bootstrap. Boca Raton: CRC Press. ISBN: 9780412042317
  5. Guttikunda, S. K., Goel, R., & Pant, P. (2014). Nature of air pollution, emission sources, and management in the Indian cities. Atmospheric Environment, 95, 501–510. doi:10.1016/j.atmosenv.2014.07.006
  6. Gurjar, B. R., Butler, T. M., Lawrence, M. G., & Lelieveld, J. (2008). Evaluation of emissions and air quality in megacities. Atmospheric Environment, 42(7), 1593–1606. doi:10.1016/j.atmosenv.2007.10.048.