VN1 forecasting competition

In 2024, I organized VN1, a forecasting competition on real retail data. In total, 978 participants registered to forecast 13 weeks of sales. Download the data and the official scoring code, and test your models against the winners.

Sponsors: Flieber, Syrup and SupChains. Platform: DataSource.ai.

  • 15,053 weekly sales series
  • 20,000€ in prizes
  • 46.4% winning score, lower is better

The task: Forecast 13 weeks of sales for 15,053 product-warehouse combinations.

Flieber, one of the sponsors, provided the sales of 46 e-vendors. Each series is one product in one warehouse. The data covers 11,171 products in 328 warehouses.

The history covers 170 weeks, from July 2020 to October 2023. Competitors also received the actual historical prices, based on transactions.

Some products ran out of stock in the past. The data does not mark these weeks, so they look like weeks without demand. We only kept products that never ran out of stock in the evaluation period.

  1. 12 September to 3 October 2024

    Phase 1, the warm-up

    Competitors forecast 13 weeks from the 170 weeks of history. A live leaderboard showed their scores.

  2. 3 to 17 October 2024

    Phase 2, the final

    We added the 13 weeks of phase 1 to the history. Competitors then forecast the next 13 weeks. Only their last forecast counted, and no score showed until the end.

  3. 13 November 2024

    The webinar

    The top 5 explained their models.

    Watch it

The score that ranked the participants

MAE% + |Bias%|

MAE% is the sum of the absolute errors, divided by total sales. Bias% is the total error, divided by total sales. A perfect forecast scores 0%.

This score offers an excellent trade-off between simplicity and business value. I explain it in Data Science for Supply Chain Forecasting and Demand Forecasting Best Practices.

The code
abs_err = np.nansum(abs(forecast - sales))
err = np.nansum(forecast - sales)
score = abs_err + abs(err)
score /= np.nansum(sales)

Few beat the naïve forecast

Phase 2 scores of the top 5 and of the two benchmarks. Lower is better.
Phase 2Score
1stJakub Figura and Philip Stubbs46.4%
2ndJustin Furlotte46.6%
3rdArsa Nikzad47.6%
4thAntoine Schwartz47.7%
5thAn Hoang48.1%
Naïve forecast50.7%
12-week moving average80.5%

Competitors had to beat a 12-week moving average, and most of them did. Because of the seasonality in the data, the naïve forecast did exceptionally well in phase 2, while the 12-week moving average did very poorly.

Usually, a naïve forecast is easy to beat, so I do not advise it as a benchmark.

Units sold per week, all 15,053 series together

The top 5 explain their models

Webinar, November 2024

What the winners did

After the competition, I interviewed the top 20 about their methods.

  • They tested many models before they chose one. Their key skill was a fast and robust way to evaluate models.
  • LightGBM was the most used model at the top.
  • Most of them combined several models, or several runs of the same model.
  • Half wrote less than 300 lines of code. Nearly all solutions ran in less than 10 minutes.
  • Only two flagged outliers. None used Facebook Prophet.

Their methods match the practices we apply on every client project.

Try it yourself

You can replay the competition on your own and score your forecasts with the official code.

  1. Replay phase 1. Train on Phase 0 and forecast 13 weeks. Score your forecast against Phase 1 sales.
  2. Replay phase 2. Add Phase 1 to the history and forecast the next 13 weeks. Score it against Phase 2 sales.
What the zip holds
Phase 0 - Sales.csvUnits sold per week, July 2020 to October 2023
Phase 0 - Price.csvPrice per week, from the transactions
Phase 1 - Sales.csv, Price.csv13 weeks, October 2023 to January 2024
Phase 2 - Sales.csvThe final ranking, January to April 2024
Top 5 forecastsThe five winning forecasts, by rank
vn1_score.pyThe official score
vn1_benchmarks.pyMoving average and naïve forecast
README.mdThe rules, the data and the score

The zip holds no phase 2 prices, because competitors never received them.

How to cite VN1. The data is free to use. When you use it, please cite the competition.

Vandeput, N. (2024). VN1 Forecasting – Accuracy Challenge. Organized by Nicolas Vandeput (SupChains). Data provided by Flieber. Sponsored by Flieber, Syrup and SupChains. https://supchains.com/vn1-forecasting-competition/

BibTeX

@misc{vandeput2024vn1,
  author       = {Vandeput, Nicolas},
  title        = {{VN1} Forecasting -- Accuracy Challenge},
  year         = {2024},
  howpublished = {\url{https://supchains.com/vn1-forecasting-competition/}},
  note         = {Organized by Nicolas Vandeput (SupChains). Data provided by Flieber. Sponsored by Flieber, Syrup and SupChains.}
}

The original competition pages stay online on DataSource.ai.

Nicolas Vandeput speaking on stage

Free chapters of our latest book

You also get the SupChains Way, then one lesson every two weeks.