Import a custom Scikit Learn script
This page describes the Machine Learning Manager, a service of the ForePaaS Legacy Platform that is not available on OVHcloud Data Platform. See the current documentation.
The Scikit Learn framework allows you to upload a Python .py file containing an estimator compatible with the Scikit Learn library. The supported libraries typically include (but aren't restricted to): Scikit Learn, XGBoost, lightgbm, ...
When you select this framework, some packages are imported by default in the pipeline's environment:
- scikit-learn
- the libraries that come out of the box with ForePaaS' SDK
Any other library used by your custom script has to be specified in the Python libraries field of the Dependencies panel (along with its version, if you wish to not use the latest one).

💡 You have the possibility to load a requirements.txt file in order to add all necessary libraries in one go.
By default, new pipelines will be using Python 3.9.9.
Upload a custom estimator
To upload a custom estimator, either drag your .py file into the Training Script box or click the box to open the file explorer.

When a Training job is launched, this .py file is the file that will be executed. It must contain a function that has the following two requirements:
- It must have
eventas its first argument - It must return a fitted estimator based on the training dataset
This function's name must be written down in the function name box.

Below is a sample code of a basic custom Scikit Learn estimator:
With the above code, you would have to specify my_random_forest as the function name.
While the previous code will not be able to use any of the configurations made in the Training and Tuning steps, you have the possibility to connect to other parts of your pipeline configuration from your estimator thanks to the ForePaaS SDK. This allows you to use all the features provided by ForePaaS pipelines with your custom estimator.
The following configurations can be integrated:
- Use a validation configuration from the Training step
- Use a validation score function from the Training step
- Use hyper-parameters specified in the Tuning step
!> Since only the custom .py script is executed during Training jobs, failing to integrate the aforementioned configurations into the Python script will result in any respective interface-originated modification to be ignored by the pipeline.
Sample code
This example combines all three aforementioned connectors:
With the above code, you would have to specify my_random_forest as the function name.

