I am using a Linear Regression classifier to predict some values. I already figured the basic part of the out and now it looks like this:
import time as ti
import pandas as pd
import numpy as np
from matplotlib import pyplot as plt
import csv
from sklearn.datasets import load_boston
from sklearn import preprocessing, svm
from sklearn.model_selection import train_test_split
from sklearn import linear_model
from scipy.interpolate import *
import datetime
data = pd.read_csv(r"C:\Users\simon\Desktop\Datenbank\visualisierung\includes\csv.csv")
x = np.array(data["day"])
y = np.array(data["balance"])
reg = linear_model.LinearRegression()
X_train, X_test, y_train, y_test, i_train, i_test = train_test_split(x, y, data.index, test_size=0.2, random_state=4)
X_train = X_train.reshape(-1, 1)
X_test = X_test.reshape(-1, 1)
i_train = i_train.values.reshape(-1, 1)
i_test = i_test.values.reshape(-1, 1)
reg.fit(i_train, y_train)
print(reg.score(i_test, y_test))
252128,6/6/19
252899,7/6/19
253670,8/6/19
254441,9/6/19
I have 27 rows of those in total.
It doesn't work for some reason.
UndefinedMetricWarning: R^2 score is not well-defined with less than two samples.
The dtypes and shapes are:
X_train, X_test = object #dtype
X_train = (21,) #shape
X_test = (6,) #shape
y_train, y_test = int64 #dtype
y_train, y_test = (1, 21) #shape
i_train, i_test = int64 #dtype
i_train, i_test = (1, 21) #shape
X_train, X_test, y_train, y_test, i_train, i_test are all a:
<class 'numpy.ndarray'>
I could imagine that thats because i dont have enough examples.
Why does this happen and how can i prevent it?
x
,X_train
etc. Shape and dtype. – hpaulji_train
to the fit function instead of the featuresX_train
? – AlCorreia