label-encoder encoding missing values

Question

I am using the label encoder to convert categorical data into numeric values. How does LabelEncoder handle missing values? Output: For the above example, label encoder changed NaN values to a category. How would I know which category represents missing values? Answer Don&#8217;t use LabelEncoder with missing …

Accepted Answer

Don&#8217;t use LabelEncoder with missing values. I don&#8217;t know which version of scikit-learn you&#8217;re using, but in 0.17.1 your code raises TypeError: unorderable types: str() > float().As you can see in the source it uses numpy.unique against the data to encode, which raises TypeError if missing values are found. If you want to encode missing values, first change its type to a string:a[pd.isnull(a)]  = 'NaN'

Advertisement

Answer