Skip to content

pandas max

Find the largest value — and find where it is.

Which is the biggest? And where is it? Those look like one question, but Pandas treats them as two, and gives you a separate method for each. .max() returns the value. .idxmax() returns the index label of the row where that value lives. .min() and .idxmin() are the same pair, pointed downward.

I think that splitting them up is exactly why idxmax stays mysterious for so long. On its own it looks like an oddly named cousin of max. Sitting next to max, it is obvious: read the name as "index of the max," and you have it. Together they also get you what you usually wanted all along — not the number, not the label, but the whole winning row:

df.loc[df['column'].idxmax()]

That one line is worth the price of admission. .idxmax() hands you a label, .loc[] takes labels, and the two click together.

Official documentation: DataFrame.max and DataFrame.idxmax — with min and idxmin alongside them.

The arguments that earn their keep

All four methods take the same three:

df.max(axis='index',        # down each column (default), or across each row
       skipna=True,         # ignore NaN, or let a single NaN swallow the answer
       numeric_only=False)  # consider every column, or only the numeric ones

df.idxmax(axis='index', skipna=True, numeric_only=False)

axis is the one people forget. The default runs down each column and gives you one answer per column. Pass axis='columns' and it runs across each row instead, so .idxmax(axis='columns') tells you, for every row, which column name holds its largest value. That is a genuinely different and very useful question.

skipna=True is the default and it is almost always what you want. Setting it to False says "a missing value means I cannot answer," and on .max() you get NaN back. On .idxmax() there is no NaN label to return, so it raises instead.

numeric_only=False is the default on a data frame, which means Pandas will happily take the maximum of your text columns too. That is the source of the first mistake below.

A worked example, on real data

Here is the Global Coal Plant Tracker, the workbook behind Bamboo Weekly #64. One row per generating unit, 13,906 of them:

import pandas as pd

url = ('https://www.bambooweekly.com/content/files/wp-content/uploads/2024/02/'
       'global-coal-plant-tracker-january-2024.xlsx')

units = pd.read_excel(url, sheet_name='Units',
                      usecols=['Country', 'Region', 'Plant name',
                               'Status', 'Capacity (MW)', 'Start year'])

operating = units.loc[lambda df_: df_['Status'] == 'operating']

What is the largest coal-burning unit running anywhere in the world?

operating['Capacity (MW)'].max()
1350.0

A number, with no idea whose it is. That is what .idxmax() is for:

operating['Capacity (MW)'].idxmax()
2575

A label — here an integer, because the frame came in with a default index, but a label all the same. Feed it to .loc[] and the answer arrives in full:

operating.loc[operating['Capacity (MW)'].idxmax()]
Country                                   China
Plant name       Huaibei Pingshan power station
Capacity (MW)                            1350.0
Status                                operating
Start year                               2022.0
Region                                     Asia
Name: 2575, dtype: object

Now let me aggregate to the country level, which gives me a frame whose index is made of names rather than numbers:

by_country = (
    operating
    .groupby('Country')['Capacity (MW)']
    .agg(total='sum', largest_unit='max', units='count')
)

by_country.head()
                          total  largest_unit  units
Country
Argentina                 495.0         375.0      2
Australia               22403.0         750.0     53
Bangladesh               4775.0         660.0     10
Bosnia and Herzegovina   2090.0         300.0     10
Botswana                  732.0         150.0      8

.idxmax() on a data frame runs down every column and returns one label each:

by_country.idxmax()
total           China
largest_unit    China
units           China
dtype: str

China three times over, which surprises nobody. The other end is more interesting — .idxmin() finds the country with the smallest operating fleet, and .loc[] shows it whole:

by_country.loc[by_country['total'].idxmin()]
total           64.0
largest_unit    32.0
units            2.0
Name: Guadeloupe, dtype: float64

Turn the frame sideways and axis='columns' starts earning its keep. Pivoting capacity by status gives one column per status, so asking which column is largest asks which phase of life each country's coal fleet is mostly in:

capacity = (
    units
    .loc[lambda df_: df_['Status'].isin(['operating', 'construction', 'retired'])]
    .pivot_table(index='Country', columns='Status',
                 values='Capacity (MW)', aggfunc='sum')
)

capacity.idxmax(axis='columns').head(8)
Country
Argentina                 operating
Australia                 operating
Austria                     retired
Bangladesh                operating
Belgium                     retired
Bosnia and Herzegovina    operating
Botswana                  operating
Brazil                    operating
dtype: str

Austria and Belgium have retired more coal capacity than they still run. The same table read the default way, down the columns, names the champion of each status:

capacity.idxmax()
Status
construction            China
operating               China
retired         United States
dtype: str

Finally, the pairing that matters most in daily work. groupby(...).max() gives you a table of numbers:

operating.groupby('Region')['Capacity (MW)'].max()
Region
Africa       800.0
Americas    1300.0
Asia        1350.0
Europe      1100.0
Oceania      750.0
Name: Capacity (MW), dtype: float64

Whereas groupby(...).idxmax() gives you a table of labels — one row label per group — which you then hand straight back to .loc[] to recover the winning rows themselves:

operating.loc[operating.groupby('Region')['Capacity (MW)'].idxmax(),
              ['Region', 'Country', 'Plant name', 'Capacity (MW)']]
         Region        Country                      Plant name  Capacity (MW)
11523    Africa   South Africa            Kusile power station          800.0
12436  Americas  United States                      Amos Plant         1300.0
2575       Asia          China  Huaibei Pingshan power station         1350.0
6922     Europe        Germany           Datteln power station         1100.0
63      Oceania      Australia       Kogan Creek power station          750.0

That is the biggest operating coal unit on each continent, named, in three lines. groupby().max() could never have told me the plant names, because by the time it has taken the maximum, the rows are gone.

Two mistakes people make

Reaching for .idxmax() when you wanted .max(), or the reverse. One returns the value, the other the label it sits at, and df.loc[…idxmax()] returns the whole row. That distinction, the ties that go silently to the first occurrence, and the all-NaN column that raises on idxmax while max returns NaN, are all on the idxmax and idxmin pages.

.max() on text compares lexically, and does it silently. With numeric_only=False as the default, a data frame of mixed types answers anyway:

units.max()
Country                       Zimbabwe
Plant name       Štavalj Power Station
Capacity (MW)                   6300.0
Status                         shelved
Start year                      2037.0
Region                         Oceania
dtype: object

Zimbabwe is not the largest country and "shelved" is not the largest status — those are just the strings that sort last, and "Štavalj" only wins because Š sorts after Z in Unicode. Nothing warns you. Pass numeric_only=True when you mean numbers.

Where it shows up in Bamboo Weekly

#9: US house prices finds the peak of the index with df['price'].max(), then asks which row it was — the question that leads straight to idxmax, and the reason the two methods are usually typed within a line of each other.

#34: House of Representatives uses groupby('state')['district'].max() to find how many districts each state has — a plain grouped maximum, where the number is the whole answer and no label is needed.

#54: Household debt averages six categories of delinquent loans by year and then compares them, which is the case for axis='columns': the reducer runs across a row rather than down a column.

Practice it

Work through a max exercise, with instant feedback and no signup required: practice.lernerpython.com/bamboo-weekly/max/

Go deeper

If you want the top few rather than the top one, that is nlargest and its twin nsmallest, which beat sorting the whole frame with sort_values. If you want several of these answers at once, agg takes them by name — .agg(['mean', 'min', 'idxmin', 'max', 'idxmax']) gives you the extremes and their locations in a single pass. And since every idxmax ends in a lookup, loc is the method to be fluent in before this one.

More Pandas videos on Python and Pandas with Reuven Lerner.

Bamboo Weekly is the practice. If you want the structured version — full courses with downloadable Jupyter notebooks, plus live Pandas office hours when you get stuck — that is LernerPython+Data. A paid Bamboo Weekly subscription is included with it.

See it on real data

Below are the 23 Bamboo Weekly exercises that use max on real-world data — try each one, then study the worked solution.

Part of the Pandas Methods Index. See also practice by skill.