This archive containts the research implementation in JAVA by Panagiotis Mandros <pmandros@mpi-inf.mpg.de> of
Panagiotis Mandros, Mario Boley, Jilles Vreeken, "Discovering Reliable Approximate Functional Dependencies". In: Proceeedings of the 23th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD'17), Halifax, Canada, 2017.

:: How To Run

The code creates as an output, a file in the -OUTPUTFOLDER folder, with the dataset name, alpha used, k for top-k and a timestamp of the run. The file contains various statistics of the search, and the top-k patterns discovered.  Every pattern contains the entropy of the target, the mutual information, the expected mutual information under the null, the reliable fraction of information score, and the uncorrected reliable fraction of information score. The code runs with .arff file format (WEKA format).

arguments:
	Obligatory
		-DATASET  (dataset filename)
		-OUTPUTFOLDER  (output folder)
	Optional
		-TARGET (the index of the target variable (starting from 0), default is the last column)
		-ALPHA   (alpha to use, default is 1)
		-K		(number for top-k, default is 1)

example run:
	For the example dataset, abalone.arff, located in this folder, the following command

		java -cp  FoBnB.jar eda.mmci.uni_saarland.de.fobnb.FoBnB -DATASET abalone.arff -OUTPUTFOLDER exampleOutput/ -K 5 -ALPHA 1 

	produces the output found in the exampleOutput folder.


:: Datasets Used

Datasets used in the Section "Optimization performance" are from the KEEL repository: http://sci2s.ugr.es/keel/datasets.php. Note that the code includes an equal-frequency discretization in 5 bins for ordinal data. This feature is activated only for ordinal data, IF any. Alternatively one can discretize the data before running the jar file.