Criteo Uplift Prediction Dataset - Criteo AI Lab

Criteo Uplift Prediction Dataset

By: Criteo AI Lab / 31 May 2018

Criteo Uplift Modeling Dataset

This dataset is released along with the paper:

“ A Large Scale Benchmark for Uplift Modeling”

Eustache Diemert, Artem Betlei, Christophe Renaudin; (Criteo AI Lab), Massih-Reza Amini (LIG, Grenoble INP)

This work was published in: AdKDD 2018 Workshop, in conjunction with KDD 2018.

When using this dataset, please cite the paper with following bibtex:

@inproceedings{Diemert2018,
author = {{Diemert Eustache, Betlei Artem} and Renaudin, Christophe and Massih-Reza, Amini},
title={A Large Scale Benchmark for Uplift Modeling},
publisher = {ACM},
booktitle = {Proceedings of the AdKDD and TargetAd Workshop, KDD, London, United Kingdom, August, 20, 2018},
year = {2018}
}

Data description

This dataset is constructed by assembling data resulting from several incrementality tests, a particular randomized trial procedure where a random part of the population is prevented from being targeted by advertising. It consists of 25M rows, each one representing a user with 11 features, a treatment indicator and 2 labels (visits and conversions).

Privacy

For privacy reasons the data has been sub-sampled non-uniformly so that the original incrementality level cannot be deduced from the dataset while preserving a realistic, challenging benchmark. Feature names have been anonymized and their values randomly projected so as to keep predictive power while making it practically impossible to recover the original features or user context.

CRITEO DATA TERMS OF USE

By exercising the Licensed Rights (defined below), You accept and agree to be bound by the terms and conditions of this Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International Public License (“Public License”). To the extent this Public License may be interpreted as a contract, You are granted the Licensed Rights in consideration of Your acceptance of these terms and conditions.

Fields

Here is a detailed description of the fields (they are comma-separated in the file):

Key figures

Tasks

The dataset was collected and prepared with uplift prediction in mind as the main task. Additionally we can foresee related usages such as but not limited to:

ERRATUM

Non uniformity of the incrementality level across advertisers caused the first version of the dataset to have a leak: uplift prediction could be artificially improved by differentiating advertisers using individual features (distribution of features being advertiser-dependent). For this reason, we release an un-biased version of the dataset containing the same fields. Here are its corresponding key figures:

Key figures  
Format: CSV  
Size: 297M (compressed)  
Rows: 13,979,592  
Average Visit Rate: .046992  
Average Conversion Rate: .00292  
Treatment Ratio: .85  

Download instructions

To download the un-biased dataset click here