Supplementary material concerning repeatability of
Causal Inference by Direction of Information

Jilles Vreeken

published in SIAM International Conference on Data Mining (SDM 2015), Vancouver, British Columbia, Canada (2015)

This page provides additional information about our experiments and assist in reproducing the results of the Ergo algorithm. We provide this in addition to our publication [ download full text PDF ] published at SDM 2015 conference.
 

Computation of Ergo Results


Assume that we want to infer if X causes Y, or Y causes X, or they are just correlated. Ergo was implemented by Hoang-Vu Nguyen and Jilles Vreeken. The code Ergo is here. A complementary JAR file is here. In order to execute them you can use the following command line structure:

Parameter Meaning
-FILE_INPUT name of input file
-FILE_RUNTIME_OUTPUT name of output file for causal scores and runtime
-NUM_ROWS number of data points
-NUM_MEASURE_COLS number of dimensions
-FIELD_DELIMITER delimiter of the input file (default is semicolon)
-ALPHA parameter epsilon the paper
-CLUMPS parameter c in the paper
-MAXDIMX maximum dimension index of X, used when X and/or Y are multivariate


When X and Y are univariate: Sample data is here. X and Y are in the same input file. X is in the first dimension and Y is in the second.

java PairsSingle -FILE_INPUT pair.txt -FILE_RUNTIME_OUTPUT rt.txt -NUM_ROWS 10226 -NUM_MEASURE_COLS 2 -ALPHA 0.3 -CLUMPS 10


When X and Y are multivariate: Sample data is here. X and Y are in the same input file. The dimension indices of X are from 0 to 5. The dimension indices of Y are from 6 to 7.

java Pairs -FILE_INPUT pair2.txt -FILE_RUNTIME_OUTPUT rt2.txt -NUM_ROWS 120 -NUM_MEASURE_COLS 8 -ALPHA 0.3 -CLUMPS 10 -MAXDIMX 5
 

Download


If you publish results based on our material, then please include a reference to our paper.