Soy nuevo tanto en Python como en Tensorflow. Estoy tratando de ejecutar el archivo del tutorial de detección de objetos desde la API de detección de objetos de Tensorflow , pero no puedo encontrar dónde puedo obtener las coordenadas de los cuadros delimitadores cuando se detectan objetos.
Código relevante:
# The following processing is only for single image detection_boxes = tf.squeeze(tensor_dict['detection_boxes'], [0]) detection_masks = tf.squeeze(tensor_dict['detection_masks'], [0])El lugar donde supongo que se dibujan los cuadros delimitadores es así:
# Visualization of the results of detection. vis_util.visualize_boxes_and_labels_on_image_array( image_np, output_dict['detection_boxes'], output_dict['detection_classes'], output_dict['detection_scores'], category_index, instance_masks=output_dict.get('detection_masks'), use_normalized_coordinates=True, line_thickness=8) plt.figure(figsize=IMAGE_SIZE) plt.imshow(image_np) Intenté imprimir output_dict['detection_boxes'] pero no estoy seguro de qué significan los números. hay muchos
array([[ 0.56213236, 0.2780568 , 0.91445708, 0.69120586], [ 0.56261235, 0.86368728, 0.59286624, 0.8893863 ], [ 0.57073039, 0.87096912, 0.61292225, 0.90354401], [ 0.51422435, 0.78449738, 0.53994244, 0.79437423], ...... [ 0.32784131, 0.5461576 , 0.36972913, 0.56903434], [ 0.03005961, 0.02714229, 0.47211722, 0.44683522], [ 0.43143299, 0.09211366, 0.58121657, 0.3509962 ]], dtype=float32)Encontré respuestas para preguntas similares, pero no tengo una variable llamada cuadros como ellos. ¿Cómo puedo obtener las coordenadas?
Intenté imprimir output_dict['detection_boxes'] pero no estoy seguro de qué significan los números
Puedes comprobar el código por ti mismo. visualize_boxes_and_labels_on_image_array se define aquí .
Tenga en cuenta que está pasando use_normalized_coordinates=True . Si rastrea las llamadas de función, verá sus números [ 0.56213236, 0.2780568 , 0.91445708, 0.69120586] etc. son los valores [ymin, xmin, ymax, xmax] donde la imagen coordina:
(left, right, top, bottom) = (xmin * im_width, xmax * im_width, ymin * im_height, ymax * im_height)son calculados por la función:
def draw_bounding_box_on_image(image, ymin, xmin, ymax, xmax, color='red', thickness=4, display_str_list=(), use_normalized_coordinates=True): """Adds a bounding box to an image. Bounding box coordinates can be specified in either absolute (pixel) or normalized coordinates by setting the use_normalized_coordinates argument. Each string in display_str_list is displayed on a separate line above the bounding box in black text on a rectangle filled with the input 'color'. If the top of the bounding box extends to the edge of the image, the strings are displayed below the bounding box. Args: image: a PIL.Image object. ymin: ymin of bounding box. xmin: xmin of bounding box. ymax: ymax of bounding box. xmax: xmax of bounding box. color: color to draw bounding box. Default is red. thickness: line thickness. Default value is 4. display_str_list: list of strings to display in box (each to be shown on its own line). use_normalized_coordinates: If True (default), treat coordinates ymin, xmin, ymax, xmax as relative to the image. Otherwise treat coordinates as absolute. """ draw = ImageDraw.Draw(image) im_width, im_height = image.size if use_normalized_coordinates: (left, right, top, bottom) = (xmin * im_width, xmax * im_width, ymin * im_height, ymax * im_height)Tengo exactamente la misma historia. Obtuve una matriz con aproximadamente cien cuadros ( output_dict['detection_boxes'] ) cuando solo se mostraba uno en una imagen. Al profundizar en el código que dibuja un rectángulo, pude extraerlo y usarlo en mi inference.py :
#so detection has happened and you've got output_dict as a # result of your inference # then assume you've got this in your inference.py in order to draw rectangles vis_util.visualize_boxes_and_labels_on_image_array( image_np, output_dict['detection_boxes'], output_dict['detection_classes'], output_dict['detection_scores'], category_index, instance_masks=output_dict.get('detection_masks'), use_normalized_coordinates=True, line_thickness=8) # This is the way I'm getting my coordinates boxes = output_dict['detection_boxes'] # get all boxes from an array max_boxes_to_draw = boxes.shape[0] # get scores to get a threshold scores = output_dict['detection_scores'] # this is set as a default but feel free to adjust it to your needs min_score_thresh=.5 # iterate over all objects found for i in range(min(max_boxes_to_draw, boxes.shape[0])): # if scores is None or scores[i] > min_score_thresh: # boxes[i] is the box which will be drawn class_name = category_index[output_dict['detection_classes'][i]]['name'] print ("This box is gonna get used", boxes[i], output_dict['detection_classes'][i])La respuesta anterior no funcionó para mí, tuve que hacer algunos cambios. Entonces, si eso no ayuda, tal vez intente esto.
# This is the way I'm getting my coordinates boxes = detections['detection_boxes'].numpy()[0] # get all boxes from an array max_boxes_to_draw = boxes.shape[0] # get scores to get a threshold scores = detections['detection_scores'].numpy()[0] # this is set as a default but feel free to adjust it to your needs min_score_thresh=.5 # # iterate over all objects found coordinates = [] for i in range(min(max_boxes_to_draw, boxes.shape[0])): if scores[i] > min_score_thresh: class_id = int(detections['detection_classes'].numpy()[0][i] + 1) coordinates.append({ "box": boxes[i], "class_name": category_index[class_id]["name"], "score": scores[i] }) print(coordinates)Aquí, cada elemento (diccionario) en la lista de coordenadas es un cuadro que se dibujará en la imagen con las coordenadas de los cuadros (normalizadas), el nombre de la clase y la puntuación.