Metadata-Version: 2.1
Name: prolothar-fdd
Version: 1.0.0
Summary: algorithms for functional dependency discovery
Home-page: https://gitlab.dillinger.de/KI/DataScience/processmining/prolothar-fdd
Author: Boris Wiegand
Author-email: boris.wiegand@stahl-holding-saar.de
License: UNKNOWN
Description: # Prolothar
        
        Algorithms for functional dependency discovery
        
        ## Getting Started
        
        These instructions will get you a copy of the project up and running on your local machine for development and testing purposes. See deployment for notes on how to deploy the project on a live system.
        
        ### Prerequisites
        
        Python 3.8+
        
        ### Installing
        
        ```
        python -m pip install
               -i http://nexus.int.shsservices.de/repository/ki-python-releases/simple
               --trusted-host nexus.int.shsservices.de prolothar-fdd
        ```
        
        If pip is already configured to use the nexus repository:
        
        ```
        pip install prolothar-fdd
        ```
        
        ## Usage
        
        ```
        from prolothar_common.models.dataset import Dataset
        from prolothar_common.models.dataset.instance import Instance
        from prolothar_fdd.fodiscovery import GreedyFoDiscovery
        
        #build your dataset. define categorical and numerical attributes
        dataset = Dataset(['color'],['size'])
        #add instances to your dataset where each instance has a unique ID
        dataset.add_instance(Instance('instance_42', {'color': 'red', 'size': 100}))
        #see test_fodiscovery for a complete example
        
        #configure the functional dependency discovery
        #larger beam width makes discovery more exact but also slower.
        #for exact discovery use from prolothar_fdd.fodiscovery import BranchAndBoundFoDiscovery
        #number_of_results controls how many of the top-k candidates will be returned
        discovery = GreedyFoDiscovery(verbose=True, beam_width=10, nr_of_results=3)
        
        #run the discovery
        functional_pattern_list = discovery.run(dataset, 'size')
        
        for functional_pattern in functional_pattern_list:
            print(functional_pattern)
        ```
        
        ## Running the tests
        
        ```
        make test
        ```
        
        ## Deployment
        
        When changes are pushed into the master branch, the project is bundled and
        uploaded to our Nexus automatically.
        
        ## Versioning
        
        We use [SemVer](http://semver.org/) for versioning.
        
        ## Authors
        
        * **Boris Wiegand** - boris.wiegand@stahl-holding-saar.de
        
        See also the list of [contributors](https://gitlab.dillinger.de/KI/DataScience/processmining/prolothar-fdd/-/graphs/master) who participated in this project.
        
        
Platform: UNKNOWN
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.8
Classifier: Topic :: Data Mining
Classifier: Topic :: Functional Dependency Analysis
Description-Content-Type: text/markdown
