Estoy intentando el código de esta página . Corrí hasta la parte LR (tf-idf) y obtuve resultados similares
Después de eso, decidí probar GridSearchCV . Mis preguntas a continuación:
1)
#lets try gridsearchcv #https://www.kaggle.com/enespolat/grid-search-with-logistic-regression from sklearn.model_selection import GridSearchCV grid={"C":np.logspace(-3,3,7), "penalty":["l2"]}# l1 lasso l2 ridge logreg=LogisticRegression(solver = 'liblinear') logreg_cv=GridSearchCV(logreg,grid,cv=3,scoring='f1') logreg_cv.fit(X_train_vectors_tfidf, y_train) print("tuned hpyerparameters :(best parameters) ",logreg_cv.best_params_) print("best score :",logreg_cv.best_score_) #tuned hpyerparameters :(best parameters) {'C': 10.0, 'penalty': 'l2'} #best score : 0.7390325593588823Luego calculé la puntuación f1 manualmente. ¿Por qué no coincide?
logreg_cv.predict_proba(X_train_vectors_tfidf)[:,1] final_prediction=np.where(logreg_cv.predict_proba(X_train_vectors_tfidf)[:,1]>=0.5,1,0) #https://www.statology.org/f1-score-in-python/ from sklearn.metrics import f1_score #calculate F1 score f1_score(y_train, final_prediction) 0.9839388145315489scoring='precision' ¿por qué da el siguiente error? No estoy claro principalmente porque tengo un conjunto de datos relativamente equilibrado (55-45%) y f1 , que requiere precision , se calcula sin ningún problema. #lets try gridsearchcv #https://www.kaggle.com/enespolat/grid-search-with-logistic-regression
from sklearn.model_selection import GridSearchCV grid={"C":np.logspace(-3,3,7), "penalty":["l2"]}# l1 lasso l2 ridge logreg=LogisticRegression(solver = 'liblinear') logreg_cv=GridSearchCV(logreg,grid,cv=3,scoring='precision') logreg_cv.fit(X_train_vectors_tfidf, y_train) print("tuned hpyerparameters :(best parameters) ",logreg_cv.best_params_) print("best score :",logreg_cv.best_score_) /usr/local/lib/python3.7/dist-packages/sklearn/metrics/_classification.py:1308: UndefinedMetricWarning: Precision is ill-defined and being set to 0.0 due to no predicted samples. Use `zero_division` parameter to control this behavior. _warn_prf(average, modifier, msg_start, len(result)) /usr/local/lib/python3.7/dist-packages/sklearn/metrics/_classification.py:1308: UndefinedMetricWarning: Precision is ill-defined and being set to 0.0 due to no predicted samples. Use `zero_division` parameter to control this behavior. _warn_prf(average, modifier, msg_start, len(result)) /usr/local/lib/python3.7/dist-packages/sklearn/metrics/_classification.py:1308: UndefinedMetricWarning: Precision is ill-defined and being set to 0.0 due to no predicted samples. Use `zero_division` parameter to control this behavior. _warn_prf(average, modifier, msg_start, len(result)) /usr/local/lib/python3.7/dist-packages/sklearn/metrics/_classification.py:1308: UndefinedMetricWarning: Precision is ill-defined and being set to 0.0 due to no predicted samples. Use `zero_division` parameter to control this behavior. _warn_prf(average, modifier, msg_start, len(result)) /usr/local/lib/python3.7/dist-packages/sklearn/metrics/_classification.py:1308: UndefinedMetricWarning: Precision is ill-defined and being set to 0.0 due to no predicted samples. Use `zero_division` parameter to control this behavior. _warn_prf(average, modifier, msg_start, len(result)) /usr/local/lib/python3.7/dist-packages/sklearn/metrics/_classification.py:1308: UndefinedMetricWarning: Precision is ill-defined and being set to 0.0 due to no predicted samples. Use `zero_division` parameter to control this behavior. _warn_prf(average, modifier, msg_start, len(result)) tuned hpyerparameters :(best parameters) {'C': 0.1, 'penalty': 'l2'} best score : 0.9474200393672962logreg_cv . Utilicé el siguiente método para recuperar las predicciones. ¿Hay una mejor manera de hacer lo mismo? logreg_cv.predict_proba(X_train_vectors_tfidf)[:,1]
###########################
############actualización 1
The best score in GridSearchCV is calculated by taking the average score from cross validation for the best estimators. That is, it is calculated from data that is held out during fitting. From what I can tell, you are calculating predicted values from the training data and calculating an F1 score on that. Since the model was trained on that data, that is why the F1 score is so much larger compared to the results in the grid search es esa la razón por la que obtengo los siguientes resultados #tuned hpyerparameters :(best parameters) {'C': 10.0, 'penalty': 'l2'} #best score : 0.7390325593588823
pero cuando lo hago manualmente obtengo f1_score(y_train, final_prediction) 0.9839388145315489
2)
Traté de sintonizar usando f1_micro como se sugiere en la respuesta a continuación. Ningún mensaje de error. Todavía no tengo claro por qué f1_micro no falla cuando falla la precision
from sklearn.model_selection import GridSearchCV grid={"C":np.logspace(-3,3,7), "penalty":["l2"], "solver":['liblinear','newton-cg'], 'class_weight':[{ 0:0.95, 1:0.05 }, { 0:0.55, 1:0.45 }, { 0:0.45, 1:0.55 },{ 0:0.05, 1:0.95 }]}# l1 lasso l2 ridge #logreg=LogisticRegression(solver = 'liblinear') logreg=LogisticRegression() logreg_cv=GridSearchCV(logreg,grid,cv=3,scoring='f1_micro') logreg_cv.fit(X_train_vectors_tfidf, y_train) tuned hpyerparameters :(best parameters) {'C': 10.0, 'class_weight': {0: 0.45, 1: 0.55}, 'penalty': 'l2', 'solver': 'newton-cg'} best score : 0.7894909688013136Termina con el error con precisión porque parte de su penalización es demasiado fuerte para este modelo, si verifica los resultados, obtiene 0 para el puntaje f1 cuando C = 0.001 y C = 0.01
res = pd.DataFrame(logreg_cv.cv_results_) res.iloc[:,res.columns.str.contains("split[0-9]_test_score|params",regex=True)] params split0_test_score split1_test_score split2_test_score 0 {'C': 0.001, 'penalty': 'l2'} 0.000000 0.000000 0.000000 1 {'C': 0.01, 'penalty': 'l2'} 0.000000 0.000000 0.000000 2 {'C': 0.1, 'penalty': 'l2'} 0.973568 0.952607 0.952174 3 {'C': 1.0, 'penalty': 'l2'} 0.863934 0.851064 0.836449 4 {'C': 10.0, 'penalty': 'l2'} 0.811634 0.769547 0.787838 5 {'C': 100.0, 'penalty': 'l2'} 0.789826 0.762162 0.773438 6 {'C': 1000.0, 'penalty': 'l2'} 0.781003 0.750000 0.763871Puedes comprobar esto:
lr = LogisticRegression(C=0.01).fit(X_train_vectors_tfidf,y_train) np.unique(lr.predict(X_train_vectors_tfidf)) array([0])Y que las probabilidades predichas se desvían hacia el intercepto:
# expected probability np.exp(lr.intercept_)/(1+np.exp(lr.intercept_)) array([0.41764462]) lr.predict_proba(X_train_vectors_tfidf) array([[0.58732636, 0.41267364], [0.57074279, 0.42925721], [0.57219143, 0.42780857], ..., [0.57215605, 0.42784395], [0.56988186, 0.43011814], [0.58966184, 0.41033816]])Para la pregunta sobre "obtener predicciones sobre los datos del tren", creo que esa es la única forma. El modelo se reajusta en todo el conjunto de entrenamiento utilizando los mejores parámetros, pero las predicciones o probabilidades pronosticadas no se almacenan. Si está buscando los valores obtenidos durante el entrenamiento/prueba, puede consultar cross_val_predict