Criteo Uplift Prediction Dataset - Criteo AI Lab
Criteo Uplift Prediction Dataset
By: Criteo AI Lab / 31 May 2018
Criteo Uplift Modeling Dataset
This dataset is released along with the paper:
“ A Large Scale Benchmark for Uplift Modeling”
Eustache Diemert, Artem Betlei, Christophe Renaudin; (Criteo AI Lab), Massih-Reza Amini (LIG, Grenoble INP)
This work was published in: AdKDD 2018 Workshop, in conjunction with KDD 2018.
When using this dataset, please cite the paper with following bibtex:
@inproceedings{Diemert2018,
author = {{Diemert Eustache, Betlei Artem} and Renaudin, Christophe and Massih-Reza, Amini},
title={A Large Scale Benchmark for Uplift Modeling},
publisher = {ACM},
booktitle = {Proceedings of the AdKDD and TargetAd Workshop, KDD, London, United Kingdom, August, 20, 2018},
year = {2018}
}
Data description
This dataset is constructed by assembling data resulting from several incrementality tests, a particular randomized trial procedure where a random part of the population is prevented from being targeted by advertising. It consists of 25M rows, each one representing a user with 11 features, a treatment indicator and 2 labels (visits and conversions).
Privacy
For privacy reasons the data has been sub-sampled non-uniformly so that the original incrementality level cannot be deduced from the dataset while preserving a realistic, challenging benchmark. Feature names have been anonymized and their values randomly projected so as to keep predictive power while making it practically impossible to recover the original features or user context.
CRITEO DATA TERMS OF USE
By exercising the Licensed Rights (defined below), You accept and agree to be bound by the terms and conditions of this Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International Public License (“Public License”). To the extent this Public License may be interpreted as a contract, You are granted the Licensed Rights in consideration of Your acceptance of these terms and conditions.
Fields
Here is a detailed description of the fields (they are comma-separated in the file):
- f0, f1, f2, f3, f4, f5, f6, f7, f8, f9, f10, f11: feature values (dense, float)
- treatment: treatment group (1 = treated, 0 = control)
- conversion: whether a conversion occurred for this user (binary, label)
- visit: whether a visit occurred for this user (binary, label)
- exposure: treatment effect, whether the user has been effectively exposed (binary)
Key figures
- Format: CSV
- Size: 459MB (compressed)
- Rows: 25,309,483
- Average Visit Rate: .04132
- Average Conversion Rate: .00229
- Treatment Ratio: .846
Tasks
The dataset was collected and prepared with uplift prediction in mind as the main task. Additionally we can foresee related usages such as but not limited to:
- benchmark for causal inference
- uplift modeling
- interactions between features and treatment
- heterogeneity of treatment
- benchmark for observational causality methods
ERRATUM
Non uniformity of the incrementality level across advertisers caused the first version of the dataset to have a leak: uplift prediction could be artificially improved by differentiating advertisers using individual features (distribution of features being advertiser-dependent). For this reason, we release an un-biased version of the dataset containing the same fields. Here are its corresponding key figures:
Key figures
Format: CSV
Size: 297M (compressed)
Rows: 13,979,592
Average Visit Rate: .046992
Average Conversion Rate: .00292
Treatment Ratio: .85
Download instructions
To download the un-biased dataset click here