Python Pandas – How to Delete Row with Max/Min Values

delete-rowpandaspython

I have dataframe:

   one   N    th
0   A      5      1   
1   Z      17     0   
2   A      16     0   
3   B      9      1   
4   B      17     0   
5   B      117    1   
6   XC     35     1   
7   C      85     0    
8   Ce     965    1

I'm looking the way to keep alternating 0101 in column three without doubling 0 or 1.
So, i want to delete row with min of values in case if i have two repeating 0 in column th and max values if i have repeating 1.

My base consis of 1000000 rows.

I expect to have dataframe like this:

   one   N    th
0   A      5      1   
1   Z      17     0   
3   B      9      1   
4   B      17     0    
6   XC     35     1   
7   C      85     0    
8   Ce     965    1

What is the fastest way to do it. I mean vectorized way. My attempts without result.

Best Answer

using a custom `groupby.idxmax`

You can swap the sign if "th" is 1 (to get the max instead of min), then set up a custom grouper (with diff or shift + cumsum) and perform a groupby.idxmax to select the rows to keep:

out = df.loc[df['N'].mul(df['th'].map({0: 1, 1: -1}))
             .groupby(df['th'].ne(df['th'].shift()).cumsum())
             .idxmax()]

Variant with a different method to swap the sign and to compute the group:

out = df.loc[df['N'].mask(df['th'].eq(1), -df['N'])
             .groupby(df['th'].diff().ne(0).cumsum())
             .idxmax()]

Output:

  one    N  th
0   A    5   1
1   Z   17   0
3   B    9   1
4   B   17   0
6  XC   35   1
7   C   85   0
8  Ce  965   1

Intermediates:

  one    N  th  swap  group max
0   A    5   1    -5      1   X
1   Z   17   0    17      2   X
2   A   16   0    16      2    
3   B    9   1    -9      3   X
4   B   17   0    17      4   X
5   B  117   1  -117      5    
6  XC   35   1   -35      5   X
7   C   85   0    85      6   X
8  Ce  965   1  -965      7   X

using boolean masks

The above code works for an arbitrary number of consecutive 0s or 1s. If you know that you only have up to 2 successive ones, you could also use boolean indexing, which should be significantly faster:

# has the value higher precedence than the next?
D = df['N'].mask(df['th'].eq(1), -df['N']).diff()

# is the th different from the previous?
G = df['th'].ne(df['th'].shift(fill_value=-1))

# rule for the bottom row
m1 = D.gt(0) | G

# rule for the top row
# same rule as above but shifted up
# D is inverted
# comparison is not strict in case of equality
m2 = ( D.le(0).shift(-1, fill_value=True)
      | G.shift(-1, fill_value=True)
     )

# keep rows of interest
out = df.loc[m1&m2]

Output:

  one    N  th
0   A    5   1
1   Z   17   0
3   B    9   1
4   B   17   0
6  XC   35   1
7   C   85   0
8  Ce  965   1

Intermediates:

  one    N  th       D      G     m1     m2  m1&m2
0   A    5   1     NaN   True   True   True   True
1   Z   17   0    22.0   True   True   True   True
2   A   16   0    -1.0  False  False   True  False
3   B    9   1   -25.0   True   True   True   True
4   B   17   0    26.0   True   True   True   True
5   B  117   1  -134.0   True   True  False  False
6  XC   35   1    82.0  False   True   True   True
7   C   85   0   120.0   True   True   True   True
8  Ce  965   1 -1050.0   True   True   True   True

More complex example with equal values:

   one    N  th       D      G     m1     m2  m1&m2
0    A    5   1     NaN   True   True   True   True
1    Z   17   0    22.0   True   True   True   True
2    A   16   0    -1.0  False  False   True  False
3    B    9   1   -25.0   True   True   True   True
4    B   17   0    26.0   True   True   True   True
5    B  117   1  -134.0   True   True  False  False
6   XC   35   1    82.0  False   True   True   True
7    C   85   0   120.0   True   True   True   True
8   Ce  965   1 -1050.0   True   True   True   True
9    u  123   0  1088.0   True   True   True   True # because of D.le(0)
10   v  123   0     0.0  False  False   True  False # because or D.gt(0)

NB. in case of equality, it is possible to select the first/second row or both or none, depending on the operator used (D.le(0), D.lt(0), D.gt(0), D.ge(0)).

timings

Although limited to maximum 2 consecutive "th", the boolean mask approach is ~4-5x faster. Timed on 1M rows:

# groupby + idxmax
96.4 ms ± 6.64 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)

# boolean masks
22.2 ms ± 1.48 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)

Related Solutions

Python Pandas DataFrame – How to Select Rows Based on Column Values

To select rows whose column value equals a scalar, some_value, use ==:

df.loc[df['column_name'] == some_value]

To select rows whose column value is in an iterable, some_values, use isin:

df.loc[df['column_name'].isin(some_values)]

Combine multiple conditions with &:

df.loc[(df['column_name'] >= A) & (df['column_name'] <= B)]

Note the parentheses. Due to Python's operator precedence rules, & binds more tightly than <= and >=. Thus, the parentheses in the last example are necessary. Without the parentheses

df['column_name'] >= A & df['column_name'] <= B

is parsed as

df['column_name'] >= (A & df['column_name']) <= B

which results in a Truth value of a Series is ambiguous error.

To select rows whose column value does not equal some_value, use !=:

df.loc[df['column_name'] != some_value]

The isin returns a boolean Series, so to select rows whose value is not in some_values, negate the boolean Series using ~:

df = df.loc[~df['column_name'].isin(some_values)] # .loc is not in-place replacement

For example,

import pandas as pd
import numpy as np
df = pd.DataFrame({'A': 'foo bar foo bar foo bar foo foo'.split(),
                   'B': 'one one two three two two one three'.split(),
                   'C': np.arange(8), 'D': np.arange(8) * 2})
print(df)
#      A      B  C   D
# 0  foo    one  0   0
# 1  bar    one  1   2
# 2  foo    two  2   4
# 3  bar  three  3   6
# 4  foo    two  4   8
# 5  bar    two  5  10
# 6  foo    one  6  12
# 7  foo  three  7  14

print(df.loc[df['A'] == 'foo'])

yields

     A      B  C   D
0  foo    one  0   0
2  foo    two  2   4
4  foo    two  4   8
6  foo    one  6  12
7  foo  three  7  14

If you have multiple values you want to include, put them in a list (or more generally, any iterable) and use isin:

print(df.loc[df['B'].isin(['one','three'])])

yields

     A      B  C   D
0  foo    one  0   0
1  bar    one  1   2
3  bar  three  3   6
6  foo    one  6  12
7  foo  three  7  14

Note, however, that if you wish to do this many times, it is more efficient to make an index first, and then use df.loc:

df = df.set_index(['B'])
print(df.loc['one'])

yields

       A  C   D
B              
one  foo  0   0
one  bar  1   2
one  foo  6  12

or, to include multiple values from the index use df.index.isin:

df.loc[df.index.isin(['one','two'])]

yields

       A  C   D
B              
one  foo  0   0
one  bar  1   2
two  foo  2   4
two  foo  4   8
two  bar  5  10
one  foo  6  12

Python – How to Delete a File or Folder

Use one of these methods:

pathlib.Path.unlink() removes a file or symbolic link.
pathlib.Path.rmdir() removes an empty directory.
shutil.rmtree() deletes a directory and all its contents.

On Python 3.3 and below, you can use these methods instead of the pathlib ones:

os.remove() removes a file.
os.unlink() removes a symbolic link.
os.rmdir() removes an empty directory.

Best Answer

using a custom groupby.idxmax

using boolean masks

timings

Related Solutions

Python Pandas DataFrame – How to Select Rows Based on Column Values

Python – How to Delete a File or Folder

Related Question

using a custom `groupby.idxmax`