Business
Jobs
  • About Us
  • Solutions
    • Job Postings
      Post your job and receive qualified candidates in 48h.
    • Candidate Assessments
      500+ technical and psychological tests, plus anti-fraud.
    • Headhunting
      Tailor-made executive search from start to finish.
    • Payroll + EOR
      Payroll dispersal and EOR across 15+ LATAM countries.
  • Pricing
  • Jobs

0

935
Views
regresión logística y GridSearchCV usando python sklearn

Estoy intentando el código de esta página . Corrí hasta la parte LR (tf-idf) y obtuve resultados similares

Después de eso, decidí probar GridSearchCV . Mis preguntas a continuación:

1)

 #lets try gridsearchcv #https://www.kaggle.com/enespolat/grid-search-with-logistic-regression from sklearn.model_selection import GridSearchCV grid={"C":np.logspace(-3,3,7), "penalty":["l2"]}# l1 lasso l2 ridge logreg=LogisticRegression(solver = 'liblinear') logreg_cv=GridSearchCV(logreg,grid,cv=3,scoring='f1') logreg_cv.fit(X_train_vectors_tfidf, y_train) print("tuned hpyerparameters :(best parameters) ",logreg_cv.best_params_) print("best score :",logreg_cv.best_score_) #tuned hpyerparameters :(best parameters) {'C': 10.0, 'penalty': 'l2'} #best score : 0.7390325593588823

Luego calculé la puntuación f1 manualmente. ¿Por qué no coincide?

 logreg_cv.predict_proba(X_train_vectors_tfidf)[:,1] final_prediction=np.where(logreg_cv.predict_proba(X_train_vectors_tfidf)[:,1]>=0.5,1,0) #https://www.statology.org/f1-score-in-python/ from sklearn.metrics import f1_score #calculate F1 score f1_score(y_train, final_prediction) 0.9839388145315489
  1. Si trato scoring='precision' ¿por qué da el siguiente error? No estoy claro principalmente porque tengo un conjunto de datos relativamente equilibrado (55-45%) y f1 , que requiere precision , se calcula sin ningún problema.

#lets try gridsearchcv #https://www.kaggle.com/enespolat/grid-search-with-logistic-regression

 from sklearn.model_selection import GridSearchCV grid={"C":np.logspace(-3,3,7), "penalty":["l2"]}# l1 lasso l2 ridge logreg=LogisticRegression(solver = 'liblinear') logreg_cv=GridSearchCV(logreg,grid,cv=3,scoring='precision') logreg_cv.fit(X_train_vectors_tfidf, y_train) print("tuned hpyerparameters :(best parameters) ",logreg_cv.best_params_) print("best score :",logreg_cv.best_score_) /usr/local/lib/python3.7/dist-packages/sklearn/metrics/_classification.py:1308: UndefinedMetricWarning: Precision is ill-defined and being set to 0.0 due to no predicted samples. Use `zero_division` parameter to control this behavior. _warn_prf(average, modifier, msg_start, len(result)) /usr/local/lib/python3.7/dist-packages/sklearn/metrics/_classification.py:1308: UndefinedMetricWarning: Precision is ill-defined and being set to 0.0 due to no predicted samples. Use `zero_division` parameter to control this behavior. _warn_prf(average, modifier, msg_start, len(result)) /usr/local/lib/python3.7/dist-packages/sklearn/metrics/_classification.py:1308: UndefinedMetricWarning: Precision is ill-defined and being set to 0.0 due to no predicted samples. Use `zero_division` parameter to control this behavior. _warn_prf(average, modifier, msg_start, len(result)) /usr/local/lib/python3.7/dist-packages/sklearn/metrics/_classification.py:1308: UndefinedMetricWarning: Precision is ill-defined and being set to 0.0 due to no predicted samples. Use `zero_division` parameter to control this behavior. _warn_prf(average, modifier, msg_start, len(result)) /usr/local/lib/python3.7/dist-packages/sklearn/metrics/_classification.py:1308: UndefinedMetricWarning: Precision is ill-defined and being set to 0.0 due to no predicted samples. Use `zero_division` parameter to control this behavior. _warn_prf(average, modifier, msg_start, len(result)) /usr/local/lib/python3.7/dist-packages/sklearn/metrics/_classification.py:1308: UndefinedMetricWarning: Precision is ill-defined and being set to 0.0 due to no predicted samples. Use `zero_division` parameter to control this behavior. _warn_prf(average, modifier, msg_start, len(result)) tuned hpyerparameters :(best parameters) {'C': 0.1, 'penalty': 'l2'} best score : 0.9474200393672962
  1. ¿Hay alguna manera más fácil de recuperar las predicciones sobre los datos del tren? ya tenemos el objeto logreg_cv . Utilicé el siguiente método para recuperar las predicciones. ¿Hay una mejor manera de hacer lo mismo?

logreg_cv.predict_proba(X_train_vectors_tfidf)[:,1]

###########################

############actualización 1

  1. Responda la pregunta 1 de arriba. En el comentario de la pregunta dice que The best score in GridSearchCV is calculated by taking the average score from cross validation for the best estimators. That is, it is calculated from data that is held out during fitting. From what I can tell, you are calculating predicted values from the training data and calculating an F1 score on that. Since the model was trained on that data, that is why the F1 score is so much larger compared to the results in the grid search

es esa la razón por la que obtengo los siguientes resultados #tuned hpyerparameters :(best parameters) {'C': 10.0, 'penalty': 'l2'} #best score : 0.7390325593588823

pero cuando lo hago manualmente obtengo f1_score(y_train, final_prediction) 0.9839388145315489

2)

Traté de sintonizar usando f1_micro como se sugiere en la respuesta a continuación. Ningún mensaje de error. Todavía no tengo claro por qué f1_micro no falla cuando falla la precision

 from sklearn.model_selection import GridSearchCV grid={"C":np.logspace(-3,3,7), "penalty":["l2"], "solver":['liblinear','newton-cg'], 'class_weight':[{ 0:0.95, 1:0.05 }, { 0:0.55, 1:0.45 }, { 0:0.45, 1:0.55 },{ 0:0.05, 1:0.95 }]}# l1 lasso l2 ridge #logreg=LogisticRegression(solver = 'liblinear') logreg=LogisticRegression() logreg_cv=GridSearchCV(logreg,grid,cv=3,scoring='f1_micro') logreg_cv.fit(X_train_vectors_tfidf, y_train) tuned hpyerparameters :(best parameters) {'C': 10.0, 'class_weight': {0: 0.45, 1: 0.55}, 'penalty': 'l2', 'solver': 'newton-cg'} best score : 0.7894909688013136
over 4 years ago · Santiago Trujillo
1 answers
Answer question

0

Termina con el error con precisión porque parte de su penalización es demasiado fuerte para este modelo, si verifica los resultados, obtiene 0 para el puntaje f1 cuando C = 0.001 y C = 0.01

 res = pd.DataFrame(logreg_cv.cv_results_) res.iloc[:,res.columns.str.contains("split[0-9]_test_score|params",regex=True)] params split0_test_score split1_test_score split2_test_score 0 {'C': 0.001, 'penalty': 'l2'} 0.000000 0.000000 0.000000 1 {'C': 0.01, 'penalty': 'l2'} 0.000000 0.000000 0.000000 2 {'C': 0.1, 'penalty': 'l2'} 0.973568 0.952607 0.952174 3 {'C': 1.0, 'penalty': 'l2'} 0.863934 0.851064 0.836449 4 {'C': 10.0, 'penalty': 'l2'} 0.811634 0.769547 0.787838 5 {'C': 100.0, 'penalty': 'l2'} 0.789826 0.762162 0.773438 6 {'C': 1000.0, 'penalty': 'l2'} 0.781003 0.750000 0.763871

Puedes comprobar esto:

 lr = LogisticRegression(C=0.01).fit(X_train_vectors_tfidf,y_train) np.unique(lr.predict(X_train_vectors_tfidf)) array([0])

Y que las probabilidades predichas se desvían hacia el intercepto:

 # expected probability np.exp(lr.intercept_)/(1+np.exp(lr.intercept_)) array([0.41764462]) lr.predict_proba(X_train_vectors_tfidf) array([[0.58732636, 0.41267364], [0.57074279, 0.42925721], [0.57219143, 0.42780857], ..., [0.57215605, 0.42784395], [0.56988186, 0.43011814], [0.58966184, 0.41033816]])

Para la pregunta sobre "obtener predicciones sobre los datos del tren", creo que esa es la única forma. El modelo se reajusta en todo el conjunto de entrenamiento utilizando los mejores parámetros, pero las predicciones o probabilidades pronosticadas no se almacenan. Si está buscando los valores obtenidos durante el entrenamiento/prueba, puede consultar cross_val_predict

over 4 years ago · Santiago Trujillo Report
Answer question
Find remote jobs

Discover the new way to find a job!

Top jobs
Top job categories
Business
Post vacancy Pricing Sales
Legal
Terms and conditions Privacy policy
© 2026 PeakU Inc. All Rights Reserved.
Andres GPT
Show me some job opportunities
There's an error!