Playground
Demos a reviewer can poke at.
These read from committed project outputs and a Supabase-backed sample of real shots. No demo renders an invented number: every value traces to a committed result or a real model output.
xG shot-map explorer
Real shots on an attacking-half pitch, coloured by value. Toggle between my CxG and StatsBomb's own xG, or a diff mode so over- and under-valued shots pop. Filter by team, match or pressure; the panel recomputes mean value and Brier calibration live.
440 shots
59 goals in this selection
Brier score on this selection (lower is better calibrated)
Actual goal rate 0.134. Both are shot-quality estimates; the point is that CxG sits next to StatsBomb's own number on the same shots, not in isolation.
Data provenance. A 440-shot sample from the CxG diagnostic model (opponent-adjusted-metrics), benchmarked against StatsBomb xG on the same shots. Open StatsBomb data, served from Supabase. Full set lives in the repo.
Model scorecard
Every model across the portfolio, one card each. Filter by domain. The benchmark is the point: each card sits its value beside the incumbent, and a metric with nothing to beat is greyed.
Opponent-Adjusted Football Metrics
LiveCxG diagnostic (sklearn logistic/Ridge/GBM)
Contextual expected goals
Behind the benchmark.
Opponent-Adjusted Football Metrics
LiveCxG diagnostic
Calibration
No incumbent to beat
Opponent-Adjusted Football Metrics
LiveCxA baseline
Contextual expected assists
No incumbent to beat
Opponent-Adjusted Football Metrics
LiveCxA diagnostic v1 (sklearn GradientBoostingClassifier, sigmoid-calibrated)
Contextual expected assists
Ahead of the benchmark.
Opponent-Adjusted Football Metrics
LiveCxA diagnostic v1
Precision among the model’s most confident actions
Ahead of the benchmark.
Contextual Football Metrics
LiveContextual GLM
Contextual expected goals
Ahead of the benchmark.
Contextual Football Metrics
LiveContextual GLM
Cross-validation
No incumbent to beat
Frame2Threat
LiveXGBoost + GRU ensemble
Possession-danger prediction
No incumbent to beat
Frame2Threat
LivePossessionGRU
Possession-danger prediction
No incumbent to beat
Frame2Threat
LiveXGBoost (event+360)
Pass-level line-breaking
No incumbent to beat
Football Market Intelligence
LiveMarket LR (Pinnacle, thr 0.05)
Betting-market flat-stake backtest
vs Break-even
Retail Growth Intelligence
LiveTwo-stage X-learner (LightGBM stage 2)
Uplift / campaign targeting
No incumbent to beat
Retail Growth Intelligence
LiveTwo-stage X-learner
Uplift / campaign targeting
Ahead of the benchmark.
Retail Growth Intelligence
LiveTwo-stage X-learner
Uplift ranking quality
No incumbent to beat
Multi-Horizon Demand Forecasting
ProfessionalDemand forecaster
Quarterly SKU demand (~7,000 SKUs)
vs Previous approach
Residual-Load Forecasting
ProfessionalResidual-load framework
Forward price-curve accuracy (far seasons)
vs Previous methodology
Residual-Load Forecasting
ProfessionalResidual-load framework
Christmas-period demand RMSE
vs Pre-fix
Network-Charge Forecasting
ProfessionalDUoS/TNUoS forecasting engine
Network-charge forecast accuracy
vs Prior internal approach
Attendance Forecasting
ProfessionalGridSearchCV Random Forest
Cardio attendance forecast (out-of-sample)
Ahead of the benchmark (lower is better).
Attendance Forecasting
ProfessionalGridSearchCV Random Forest
Holistic attendance forecast (out-of-sample)
Ahead of the benchmark (lower is better).
Every value traces to a committed result. A card with no incumbent is greyed: a metric with nothing to beat is not a win. MAE is lower-better and reported as a count, not a percentage.
Data provenance. Hand-typed from committed results (content/metrics.ts). Football and retail are open or synthetic data; energy and consulting are professional results.
Forecast error visualiser
Before/after accuracy across three roles. E.ON and Manor Park are professional results; UoB is the MSc capstone. Aggregate deltas only, no invented time series.
Residual RMSE improved ~0.3 to 0.5 GWh/h; Christmas fix ~0.35 GWh/h. Correlation ~0.84.
Data provenance. Aggregate figures from MASTER_PROFILE.md, generated into data/forecast.json. E.ON and Manor Park report accuracy improvement; UoB reports out-of-sample MAE, which is a count and lower-better.
Uplift decile explorer
The Qini curve and per-decile uplift from the retail X-learner. Slide to target the top K% of customers and watch the captured incremental response recompute.
+0.0444
Overall ATE
response-rate lift
+0.0754
Top decile
~1.7x the average
744.8
Qini area
vs random targeting
0.406
Spearman rank
predicted vs observed
Observed uplift by decile
Percentage-point lift in response rate, treatment vs control, within each predicted decile.
Qini curve
Cumulative incremental responders captured (pink) against random targeting (cyan).
31.1%
of total incremental response captured
236
incremental responders (model units)
1.55x
the random-targeting baseline
| Decile | Customers | Observed uplift (pp) | Cumulative captured responders |
|---|---|---|---|
| 1 | 2,052 | 7.54 | 126 |
| 2 | 2,052 | 6.36 | 236 |
| 3 | 2,052 | 4.83 | 318 |
| 4 | 2,052 | 1.47 | 344 |
| 5 | 2,052 | 2.30 | 384 |
| 6 | 2,051 | 5.66 | 481 |
| 7 | 2,052 | 5.32 | 574 |
| 8 | 2,052 | 5.40 | 667 |
| 9 | 2,052 | 3.03 | 717 |
| 10 | 2,052 | 2.64 | 760 |
Synthetic retail data; demonstrates method, not a measured commercial outcome. The Qini series is derived from committed per-decile counts; headline figures are read straight from the committed model-comparison output.
Data provenance. Per-decile rows from retail-intelligence/outputs/phase_uplift_v2_decile_summary.csv and headline figures from phase_uplift_v2_model_comparison.csv, generated into data/uplift-deciles.json. Synthetic retail data; demonstrates method, not a measured commercial outcome.