pandas, scikit-learn, xgboost and seaborn integration

Last update: Dec 30, 2022

Related tags

Machine Learning pandas-ml

Overview

pandas-ml

https://travis-ci.org/pandas-ml/pandas-ml.svg?branch=master

Overview

pandas, scikit-learn and xgboost integration.

Installation

$ pip install pandas_ml

Documentation

http://pandas-ml.readthedocs.org/en/stable/

Example

>>> import pandas_ml as pdml
>>> import sklearn.datasets as datasets

# create ModelFrame instance from sklearn.datasets
>>> df = pdml.ModelFrame(datasets.load_digits())
>>> type(df)
<class 'pandas_ml.core.frame.ModelFrame'>

# binarize data (features), not touching target
>>> df.data = df.data.preprocessing.binarize()
>>> df.head()
   .target  0  1  2  3  4  5  6  7  8 ...  54  55  56  57  58  59  60  61  62  63
0        0  0  0  1  1  1  1  0  0  0 ...   0   0   0   0   1   1   1   0   0   0
1        1  0  0  0  1  1  1  0  0  0 ...   0   0   0   0   0   1   1   1   0   0
2        2  0  0  0  1  1  1  0  0  0 ...   1   0   0   0   0   1   1   1   1   0
3        3  0  0  1  1  1  1  0  0  0 ...   1   0   0   0   1   1   1   1   0   0
4        4  0  0  0  1  1  0  0  0  0 ...   0   0   0   0   0   1   1   1   0   0
[5 rows x 65 columns]

# split to training and test data
>>> train_df, test_df = df.model_selection.train_test_split()

# create estimator (accessor is mapped to sklearn namespace)
>>> estimator = df.svm.LinearSVC()

# fit to training data
>>> train_df.fit(estimator)

# predict test data
>>> test_df.predict(estimator)
0     4
1     2
2     7
...
448    5
449    8
Length: 450, dtype: int64

# Evaluate the result
>>> test_df.metrics.confusion_matrix()
Predicted   0   1   2   3   4   5   6   7   8   9
Target
0          52   0   0   0   0   0   0   0   0   0
1           0  37   1   0   0   1   0   0   3   3
2           0   2  48   1   0   0   0   1   1   0
3           1   1   0  44   0   1   0   0   3   1
4           1   0   0   0  43   0   1   0   0   0
5           0   1   0   0   0  39   0   0   0   0
6           0   1   0   0   1   0  35   0   0   0
7           0   0   0   0   2   0   0  42   1   0
8           0   2   1   0   1   0   0   0  33   1
9           0   2   1   2   0   0   0   0   1  38

Supported Packages

scikit-learn
patsy
xgboost

Comments

Fixed imports of deprecated modules which were removed in pandas 0.24.0

Certain functions were deprecated in a previous version of pandas and moved to a different module (see #117). This PR fixes the imports of those functions.

opened by kristofve 8
REL: v0.4.0
[x] Compat/test for sklearn 0.18.0 (#81)

[x] initial fix (#81)

[x] wrapper for cross validation classes (re-enable skipped tests) (#85)

[x] tests for multioutput (#86)

[x] Update doc

[x] Compat/test for pandas 0.19.0 (#83)

[x] Update release note (#88)
opened by sinhrks 4
Importation error

I tried to import pandas_ml but it gave the error :

AttributeError: type object 'NDFrame' has no attribute 'groupby'

I'm running python3.8.1 and I installed pandas_ml via pip (version 20.0.2)

I dig in the code, error is l.80 of file series.py

@Appender(pd.core.generic.NDFrame.groupby.__doc__)

Here pandas is imported at the top of the file with a classic import pandas as pd

I guess there is a problem with the versions...

Thanks in advance for any help

opened by ierezell 2
Confusion Matrix no accessible

Hi,

I've been using confusion_matrix since it was an independent package. I've installed pandas_ml to continue using the package, but it seems that the setup.py script does not install the package.

Could it be an issue with the find_packages function?

opened by mmartinortiz 2

Seaborn Scatterplot matrix / pairplot integration

import seaborn as sns
sns.set()

df = sns.load_dataset("iris")
sns.pairplot(df, hue="species")

displays

iris_scatter_matrix

but pairplot doesn't work the same way with ModelFrame

import pandas as pd
pd.set_option('max_rows', 10)
import sklearn.datasets as datasets
import pandas_ml as pdml  # https://github.com/pandas-ml/pandas-ml
import seaborn as sns
import matplotlib.pyplot as plt
df = pdml.ModelFrame(datasets.load_iris())
sns.pairplot(df, hue=".target")

iris_modelframe

There is some useless subplots

opened by scls19fr 2

Error while running train.py from speech commands in tensorflow examples.

Have the following error: File "train.py", line 27, in <module> from callbacks import ConfusionMatrixCallback File "/home/tesseract/ayush_workspace/NLP/WakeWord/tensorflow_trainer/ml/callbacks.py", line 21, in <module> from pandas_ml import ConfusionMatrix File "/home/tesseract/anaconda3/envs/ciao/lib/python3.6/site-packages/pandas_ml/__init__.py", line 3, in <module> from pandas_ml.core import ModelFrame, ModelSeries # noqa File "/home/tesseract/anaconda3/envs/ciao/lib/python3.6/site-packages/pandas_ml/core/__init__.py", line 3, in <module> from pandas_ml.core.frame import ModelFrame # noqa File "/home/tesseract/anaconda3/envs/ciao/lib/python3.6/site-packages/pandas_ml/core/frame.py", line 18, in <module> from pandas_ml.core.series import ModelSeries File "/home/tesseract/anaconda3/envs/ciao/lib/python3.6/site-packages/pandas_ml/core/series.py", line 11, in <module> class ModelSeries(ModelTransformer, pd.Series): File "/home/tesseract/anaconda3/envs/ciao/lib/python3.6/site-packages/pandas_ml/core/series.py", line 80, in ModelSeries @Appender(pd.core.generic.NDFrame.groupby.__doc__) AttributeError: type object 'NDFrame' has no attribute 'groupby' Happening with both version 5 and 6.1

opened by ayush7 1
error for example https://pandas-ml.readthedocs.io/en/latest/xgboost.html

code from example https://pandas-ml.readthedocs.io/en/latest/xgboost.html '''import pandas_ml as pdml import sklearn.datasets as datasets df = pdml.ModelFrame(datasets.load_digits()) train_df, test_df = df.cross_validation.train_test_split() estimator = df.xgboost.XGBClassifier() train_df.fit(estimator) predicted = test_df.predict(estimator) q=1 test_df.metrics.confusion_matrix() train_df.xgboost.plot_importance()

tuned_parameters = [{'max_depth': [3, 4]}] cv = df.grid_search.GridSearchCV(df.xgb.XGBClassifier(), tuned_parameters, cv=5)

df.fit(cv) df.grid_search.describe(cv) q=1

'''

gives error ''' File "E:\Pandas\my_code\S_pandas_ml_feb27.py", line 10, in train_df.xgboost.plot_importance() File "C:\Users\sndr\Anaconda3\Lib\site-packages\pandas_ml\xgboost\base.py", line 61, in plot_importance return xgb.plot_importance(self._df.estimator.booster(),

builtins.TypeError: 'str' object is not callable ''' I use Windows and 3.6.4 |Anaconda, Inc.| (default, Jan 16 2018, 10:22:32) [MSC v.1900 64 bit (AMD64)] Python Type "help", "copyright", "credits" or "license" for more information.

opened by Sandy4321 1
pandas 0.24.0 has deprecated pandas.util.decorators

See https://pandas.pydata.org/pandas-docs/stable/whatsnew/v0.24.0.html#deprecations

This causes the import statement in https://github.com/pandas-ml/pandas-ml/blob/master/pandas_ml/core/frame.py to break.

Looks like just need to change it to 'from pandas.utils'

opened by usul83 1
'mean_absoloute_error

from sklearn import metrics print('MAE:',metrics.mean_absoloute_error(y_test,y_pred)) module 'sklearn.metrics' has no attribute 'mean_absoloute_error This error is occurred..any solution

opened by vikramk1507 0
AttributeError: type object 'NDFrame' has no attribute 'groupby'

AttributeError: type object 'NDFrame' has no attribute 'groupby'

from pandas_ml import ConfusionMatrix cm = ConfusionMatrix(actu, pred) cm.print_stats()

AttributeError Traceback (most recent call last) in ----> 1 from pandas_ml import confusion_matrix 2 3 cm = ConfusionMatrix(actu, pred) 4 cm.print_stats()

/usr/local/lib/python3.8/site-packages/pandas_ml/init.py in 1 #!/usr/bin/env python 2 ----> 3 from pandas_ml.core import ModelFrame, ModelSeries # noqa 4 from pandas_ml.tools import info # noqa 5 from pandas_ml.version import version as version # noqa

/usr/local/lib/python3.8/site-packages/pandas_ml/core/init.py in 1 #!/usr/bin/env python 2 ----> 3 from pandas_ml.core.frame import ModelFrame # noqa 4 from pandas_ml.core.series import ModelSeries # noqa

/usr/local/lib/python3.8/site-packages/pandas_ml/core/frame.py in 16 from pandas_ml.core.accessor import _AccessorMethods 17 from pandas_ml.core.generic import ModelPredictor, _shared_docs ---> 18 from pandas_ml.core.series import ModelSeries 19 20

/usr/local/lib/python3.8/site-packages/pandas_ml/core/series.py in 9 10 ---> 11 class ModelSeries(ModelTransformer, pd.Series): 12 """ 13 Wrapper for pandas.Series to support sklearn.preprocessing

/usr/local/lib/python3.8/site-packages/pandas_ml/core/series.py in ModelSeries() 78 return df 79 ---> 80 @Appender(pd.core.generic.NDFrame.groupby.doc) 81 def groupby(self, by=None, axis=0, level=None, as_index=True, sort=True, 82 group_keys=True, squeeze=False):

AttributeError: type object 'NDFrame' has no attribute 'groupby'

opened by gfranco008 5
AttributeError: module 'sklearn.metrics' has no attribute 'jaccard_similarity_score'

I am using scikit-learn version 0.23.1 and I get the following error: AttributeError: module 'sklearn.metrics' has no attribute 'jaccard_similarity_score' when calling the function ConfusionMatrix.

opened by petraknovak 11
Error while running train.py from speech commands in tensorflow examples. AttributeError: type object 'NDFrame' has no attribute 'groupby'

Have the following error: File "train.py", line 27, in <module> from callbacks import ConfusionMatrixCallback File "/home/tesseract/ayush_workspace/NLP/WakeWord/tensorflow_trainer/ml/callbacks.py", line 21, in <module> from pandas_ml import ConfusionMatrix File "/home/tesseract/anaconda3/envs/ciao/lib/python3.6/site-packages/pandas_ml/__init__.py", line 3, in <module> from pandas_ml.core import ModelFrame, ModelSeries # noqa File "/home/tesseract/anaconda3/envs/ciao/lib/python3.6/site-packages/pandas_ml/core/__init__.py", line 3, in <module> from pandas_ml.core.frame import ModelFrame # noqa File "/home/tesseract/anaconda3/envs/ciao/lib/python3.6/site-packages/pandas_ml/core/frame.py", line 18, in <module> from pandas_ml.core.series import ModelSeries File "/home/tesseract/anaconda3/envs/ciao/lib/python3.6/site-packages/pandas_ml/core/series.py", line 11, in <module> class ModelSeries(ModelTransformer, pd.Series): File "/home/tesseract/anaconda3/envs/ciao/lib/python3.6/site-packages/pandas_ml/core/series.py", line 80, in ModelSeries @Appender(pd.core.generic.NDFrame.groupby.__doc__) AttributeError: type object 'NDFrame' has no attribute 'groupby' Happening with both version 5 and 6.1

opened by ayush7 3

Pandas 1.0.0rc0/0.6.1 module 'sklearn.preprocessing' has no attribute 'Imputer'

SKLEARN

sklearn.preprocessing.Imputer Warning DEPRECATED

class sklearn.preprocessing.Imputer(*args, **kwargs)[source] Imputation transformer for completing missing values.

Releases(v0.6.1)

v0.6.1(Mar 5, 2019)

Source code(tar.gz)
Source code(zip)
v0.6.0(Jan 15, 2019)

Source code(tar.gz)
Source code(zip)
v0.5.0(Nov 16, 2017)

Source code(tar.gz)
Source code(zip)
v0.4.0(Oct 15, 2016)
Support scikit-learn v0.17.x and v0.18.0.

Support imbalanced-learn via .imbalance accessor.

Added pandas_ml.ConfusionMatrix class for easier classification results evaluation.

Source code(tar.gz)
Source code(zip)
v0.3.0(Oct 22, 2015)

Source code(tar.gz)
Source code(zip)
v0.2.0(Sep 12, 2015)

Source code(tar.gz)
Source code(zip)
pandas_ml-0.2.0.tar.gz(41.68 KB)
v0.1.1(Mar 13, 2015)

Source code(tar.gz)
Source code(zip)
v0.1.0(Mar 7, 2015)

Source code(tar.gz)
Source code(zip)
v0.0.1(Mar 1, 2015)

Source code(tar.gz)
Source code(zip)

Owner

GitHub Repository

Conducted ANOVA and Logistic regression analysis using matplot library to visualize the result.

Intro-to-Data-Science Conducted ANOVA and Logistic regression analysis. Project ANOVA The main aim of this project is to perform One-Way ANOVA analysi

1 Feb 06, 2022

A toolbox to iNNvestigate neural networks' predictions!

iNNvestigate neural networks! Table of contents Introduction Installation Usage and Examples More documentation Contributing Releases Introduction In

1.1k Jan 05, 2023

A Software Framework for Neuromorphic Computing

338 Dec 26, 2022

Probabilistic programming framework that facilitates objective model selection for time-varying parameter models.

Time series analysis today is an important cornerstone of quantitative science in many disciplines, including natural and life sciences as well as eco

129 Dec 24, 2022

Pandas Machine Learning and Quant Finance Library Collection

148 Dec 07, 2022

A data preprocessing package for time series data. Design for machine learning and deep learning.

152 Jan 07, 2023

Model Agnostic Confidence Estimator (MACEST) - A Python library for calibrating Machine Learning models' confidence scores

95 Dec 28, 2022

Predicting Baseball Metric Clusters: Clustering Application in Python Using scikit-learn

Clustering Clustering Application in Python Using scikit-learn This repository contains the prediction of baseball metric clusters using MLB Statcast

2 Apr 18, 2022

2D fluid simulation implementation of Jos Stam paper on real-time fuild dynamics, including some suggested extensions.

Fluid Simulation Usage Download this repo and store it in your computer. Open a terminal and go to the root directory of this folder. Make sure you ha

5 Dec 02, 2022

Code Repository for Machine Learning with PyTorch and Scikit-Learn

1.4k Jan 03, 2023

Customers Segmentation with RFM Scores and K-means

Customer Segmentation with RFM Scores and K-means RFM Segmentation table: K-Means Clustering: Business Problem Rule-based customer segmentation machin

5 Aug 10, 2022

Distributed Deep learning with Keras & Spark

Elephas: Distributed Deep Learning with Keras & Spark Elephas is an extension of Keras, which allows you to run distributed deep learning models at sc

1.6k Dec 29, 2022

Combines Bayesian analyses from many datasets.

PosteriorStacker Combines Bayesian analyses from many datasets. Introduction Method Tutorial Output plot and files Introduction Fitting a model to a d

19 Feb 13, 2022

BigDL: Distributed Deep Learning Framework for Apache Spark

BigDL: Distributed Deep Learning on Apache Spark What is BigDL? BigDL is a distributed deep learning library for Apache Spark; with BigDL, users can w

4.1k Jan 09, 2023

This repository demonstrates the usage of hover to understand and supervise a machine learning task.

Hover Example Apps (works out-of-the-box on Binder) This repository demonstrates the usage of hover to understand and supervise a machine learning tas

43 Dec 03, 2021

OptaPy is an AI constraint solver for Python to optimize planning and scheduling problems.

OptaPy is an AI constraint solver for Python to optimize the Vehicle Routing Problem, Employee Rostering, Maintenance Scheduling, Task Assignment, School Timetabling, Cloud Optimization, Conference S

208 Dec 27, 2022

Penguins species predictor app is used to classify penguins species created using python's scikit-learn, fastapi, numpy and joblib packages.

Penguins Classification App Penguins species predictor app is used to classify penguins species using their island, sex, bill length (mm), bill depth

3 Apr 05, 2022

SageMaker Python SDK is an open source library for training and deploying machine learning models on Amazon SageMaker.

SageMaker Python SDK SageMaker Python SDK is an open source library for training and deploying machine learning models on Amazon SageMaker. With the S

1.8k Jan 01, 2023

Neighbourhood Retrieval (Nearest Neighbours) with Distance Correlation.

Neighbourhood Retrieval with Distance Correlation Assign Pseudo class labels to datapoints in the latent space. NNDC is a slim wrapper around FAISS. N

1 Jan 16, 2022

Simulate & classify transient absorption spectroscopy (TAS) spectral features for bulk semiconducting materials (Post-DFT)

PyTASER PyTASER is a Python (3.9+) library and set of command-line tools for classifying spectral features in bulk materials, post-DFT. The goal of th

4 Dec 27, 2022

pandas, scikit-learn, xgboost and seaborn integration

Related tags

Overview

pandas-ml

Overview

Installation

Documentation

Example

Supported Packages

Comments

Releases(v0.6.1)

v0.6.1(Mar 5, 2019)

v0.6.0(Jan 15, 2019)

v0.5.0(Nov 16, 2017)

v0.4.0(Oct 15, 2016)

v0.3.0(Oct 22, 2015)

v0.2.0(Sep 12, 2015)

v0.1.1(Mar 13, 2015)

v0.1.0(Mar 7, 2015)

v0.0.1(Mar 1, 2015)

Owner

Conducted ANOVA and Logistic regression analysis using matplot library to visualize the result.

A toolbox to iNNvestigate neural networks' predictions!

A Software Framework for Neuromorphic Computing

Probabilistic programming framework that facilitates objective model selection for time-varying parameter models.

Pandas Machine Learning and Quant Finance Library Collection

A data preprocessing package for time series data. Design for machine learning and deep learning.

Model Agnostic Confidence Estimator (MACEST) - A Python library for calibrating Machine Learning models' confidence scores

Predicting Baseball Metric Clusters: Clustering Application in Python Using scikit-learn

2D fluid simulation implementation of Jos Stam paper on real-time fuild dynamics, including some suggested extensions.

Code Repository for Machine Learning with PyTorch and Scikit-Learn

Customers Segmentation with RFM Scores and K-means

Distributed Deep learning with Keras & Spark

Combines Bayesian analyses from many datasets.

BigDL: Distributed Deep Learning Framework for Apache Spark

This repository demonstrates the usage of hover to understand and supervise a machine learning task.

OptaPy is an AI constraint solver for Python to optimize planning and scheduling problems.

Penguins species predictor app is used to classify penguins species created using python's scikit-learn, fastapi, numpy and joblib packages.

SageMaker Python SDK is an open source library for training and deploying machine learning models on Amazon SageMaker.

Neighbourhood Retrieval (Nearest Neighbours) with Distance Correlation.

Simulate & classify transient absorption spectroscopy (TAS) spectral features for bulk semiconducting materials (Post-DFT)