Research Dashboard · 2013–2023

Pollution kills.
Now we can prove it.

A decade of air quality data across 31 Indian states, matched against disease mortality records from the Global Burden of Disease dataset — revealing how particulate pollution drives death at a measurable, statistically significant scale.

453
Monitoring Stations
31
States Analysed
21
Disease Categories
14yr
Time Horizon
Pollution Exposure
Biological Damage
Disease Onset
Mortality
ML Forecast
Section 02
PM2.5 ↔ Disease Correlation
Pearson correlation between annual PM2.5 levels and disease death rates across 96 state-year observations. A value of +1 means pollution and deaths rise together perfectly; −1 means the opposite; 0 means no linear relationship.
Evidence Key
Direct causal link — WHO/IARC confirmed
Strong association — mechanism understood
Indirect / systemic — pollution worsens exposure
Positive Correlations — More Pollution, More Deaths
These diseases rise together with PM2.5 across states. Wider bars = stronger link.
Negative Correlations — Inverse Relationship
These diseases show lower measured death rates in high-pollution states — see note below.
Why the negative values?
Wealthier, higher-pollution urban states (Delhi, Maharashtra) have better hospital infrastructure, skewing the measured death rate downward despite higher actual disease burden. This is a known confound in ecological-level studies — not evidence that pollution protects against cardiovascular disease.
Distribution of Correlation Strengths — All 21 Diseases
Red = significant positive (p < 0.05) · Blue = significant negative · Faded = not statistically significant at 5% level
Section 04
Machine Learning Results
Two predictive tasks: forecasting future PM2.5 levels (Task A), and predicting disease death rates from pollution history (Task B). All metrics are from the real model evaluation — see Research page for full methodology.
Task A — PM2.5 Prediction
Forecasting Future Pollution
A Random Forest model was trained on monthly PM2.5 lag features, temperature, humidity, wind speed, rainfall, and seasonal cycles. Training data: 2015–2020. Test data: 2021+ — 1,202 city-month records never seen during training.
0.594
Test R²
±16.04
Avg Error (μg/m³)
RF
Model Used
1,203
Training Samples
Task B — Mortality Prediction
Predicting Deaths from Pollution History
Linear regression linking state-level PM2.5 exposure to annual disease death rates (per 100,000 population). Rate used instead of raw count to control for population size differences between states.
0.710
Best R² (Neonatal)
0.446
Cardiovascular R²
OLS
Model Used
96
State-Year Obs.
Task A — PM2.5 Trend: Training vs Test Period
Blue points = training data (2015–2020). Red points = test data (2021–2023). The model was never shown red-period data during training.
Task B — Disease Prediction Strength (R² Score)
R² = how much of the variation in disease deaths is explained by PM2.5 exposure alone. Higher = stronger predictive relationship.
Neonatal encephalopathy (R² 0.71) is the strongest predictor — severe air pollution sharply reduces birth outcomes and newborn brain oxygenation. Cardiovascular R² 0.45 means pollution explains nearly half the state-level variation in heart disease deaths.
Task A — Feature Inputs to the Random Forest Model
The model uses 9 input features for each city-month prediction. Lag features capture how last month's pollution predicts this month's. Seasonal cycles are encoded as sine and cosine of the month.