Empresas
Empregos
  • Sobre nós
  • Soluções
    • Publicação de vagas
      Publique sua vaga e receba candidatos qualificados em 48h.
    • Avaliações de candidatos
      Mais de 500 testes técnicos e psicológicos, mais anti-fraude.
    • Headhunting
      Busca executiva personalizada do início ao fim.
    • Folha de Pagamento + EOR
      Dispersão de folha e EOR em mais de 15 países da LATAM.
  • Preços
  • Empregos

0

423
Visualizações
How to calculate distance for every row in a pandas dataframe from a single point efficiently?

I have a point

point = np.array([0.07852388, 0.60007135, 0.92925712, 0.62700219, 0.16943809,
       0.34235233])

And a pandas dataframe

           a           b           c           d           e           f
0   0.025641    0.554686    0.988809    0.176905    0.050028    0.333333
1   0.027151    0.520914    0.985590    0.409572    0.163980    0.424242
2   0.028788    0.478810    0.970480    0.288557    0.095053    0.939394
3   0.018692    0.450573    0.985910    0.178048    0.118399    0.484848
4   0.023256    0.787253    0.865287    0.217591    0.205670    0.303030

I would like to calculate the distance of every row in the pandas dataframe, to that specific point

I tried

import numpy as np
d_all = list()
for index, row in df_scaled[cols_list].iterrows():
        d = np.linalg.norm(centroid-np.array(list(row[cols_list])))
        d_all += [d]
df_scaled['distance_cluster'] = d_all

My solution is really slow though (taking into account that I want to calculate the distance from other points as well.

Is there a way to do my calculations more efficiently ?

over 4 years ago · Hanz Gallego
4 Respostas
Responde à pergunta

0

You can compute vectorized Euclidean distance (L2 norm) using the formula

sqrt((a1 - b1)2 + (a2 - b2)2 + ...)

df.sub(point, axis=1).pow(2).sum(axis=1).pow(.5)

0    0.474690
1    0.257080
2    0.703857
3    0.503596
4    0.461151
dtype: float64

Which gives the same output as your current code.


Or, using linalg.norm:

np.linalg.norm(df.to_numpy() - point, axis=1)
# array([0.47468985, 0.25707985, 0.70385676, 0.5035961 , 0.46115096])
over 4 years ago · Hanz Gallego Relatório

0

Another option is use cdist which is a bit faster:

from scipy.spatial.distance import cdist
cdist(point[None,], df.values)

Output:

array([[0.47468985, 0.25707985, 0.70385676, 0.5035961 , 0.46115096]])

Some comparison with 100k rows:

%%timeit -n 10
cdist([point], df.values)
645 µs ± 36.4 µs per loop (mean ± std. dev. of 7 runs, 10 loops each)

%%timeit -n 10
np.linalg.norm(df.to_numpy() - point, axis=1)
5.16 ms ± 227 µs per loop (mean ± std. dev. of 7 runs, 10 loops each)

%%timeit -n 10
df.sub(point, axis=1).pow(2).sum(axis=1).pow(.5)
16.8 ms ± 444 µs per loop (mean ± std. dev. of 7 runs, 10 loops each)
over 4 years ago · Hanz Gallego Relatório

0

Let us do scipy

from scipy.spatial import distance
ary = distance.cdist(df.values, np.array([point]), metric='euclidean')
ary
Out[57]: 
array([[0.47468985],
       [0.25707985],
       [0.70385676],
       [0.5035961 ],
       [0.46115096]])
over 4 years ago · Hanz Gallego Relatório

0

A bit late, but you can apply the np.ligalg.norm function to the dataframe.

df['distance_cluster'] = df.apply(lambda x : np.linalg.norm(x-point),1)

Output:

#print(df['distance_cluster'])

0    0.474690
1    0.257080
2    0.703857
3    0.503596
4    0.461151
dtype: float64

However, it would be considerably slower compared to numpy solutions.

over 4 years ago · Hanz Gallego Relatório
Responde à pergunta
Encontrar trabalhos remotos

Descubra a nova forma de encontrar um emprego!

melhores empregos
Principais categorias de trabalho
Empresas
Postar vaga Preços Comercial
Jurídico
Termos e Condições Política de privacidade
© 2026 PeakU Inc. All Rights Reserved.
Andres GPT
Recomende algumas ofertas para mim
Preciso de ajuda