I want to replace some values in a column of a dataframe using a dictionary that maps the old codes to the new codes.
di = dict( { "myVar": {11:0, 204:11} } )
mydata.replace( to_replace = di, inplace = True )
But some of the new codes and old codes overlap. When using the .replace method of the dataframe I encounter the error 'Replacement not allowed with overlapping keys and values'
My current workaround is to replace replace the offending keys manually and then apply the dictionary to the remaining non-overlapping cases.
mydata.loc[ mydata.myVar == 11, "myVar" ] = 0
di = dict( { "myVar": {204:11} } )
mydata.replace( to_replace = di, inplace = True )
Is there a more compact way to do this?
I found an answer here that uses the .map method on a series in conjunction with a dictionary. Here's an example recoding dictionary with overlapping keys and values.
import pandas as pd
>>> df = pd.DataFrame( [1,2,3,4,1], columns = ['Var'] )
>>> df
Var
0 1
1 2
2 3
3 4
4 1
>>> dict = {1:2, 2:3, 3:1, 4:3}
>>> df.Var.map( dict )
0 2
1 3
2 1
3 3
4 2
Name: Var, dtype: int64
UPDATE:
With map, every value in the original series must be mapped to a new value. If the mapping dictionary does not contain all the values of the original column, the unmapped values are mapped to NaN.
>>> df = pd.DataFrame( [1,2,3,4,1], columns = ['Var'] )
>>> dict = {1:2, 2:3, 3:1}
>>> df.Var.map( dict )
0 2.0
1 3.0
2 1.0
3 NaN
4 2.0
Name: Var, dtype: float64