Find the largest value — and find where it is.
Which is the biggest? And where is it? Those look like one question, but Pandas treats them as two, and gives you a separate method for each. .max() returns the value. .idxmax() returns the index label of the row where that value lives. .min() and .idxmin() are the same pair, pointed downward.
I think that splitting them up is exactly why idxmax stays mysterious for so long. On its own it looks like an oddly named cousin of max. Sitting next to max, it is obvious: read the name as "index of the max," and you have it. Together they also get you what you usually wanted all along — not the number, not the label, but the whole winning row:
df.loc[df['column'].idxmax()]
That one line is worth the price of admission. .idxmax() hands you a label, .loc[] takes labels, and the two click together.
Official documentation: DataFrame.max and DataFrame.idxmax — with min and idxmin alongside them.
The arguments that earn their keep
All four methods take the same three:
df.max(axis='index', # down each column (default), or across each row
skipna=True, # ignore NaN, or let a single NaN swallow the answer
numeric_only=False) # consider every column, or only the numeric ones
df.idxmax(axis='index', skipna=True, numeric_only=False)
axis is the one people forget. The default runs down each column and gives you one answer per column. Pass axis='columns' and it runs across each row instead, so .idxmax(axis='columns') tells you, for every row, which column name holds its largest value. That is a genuinely different and very useful question.
skipna=True is the default and it is almost always what you want. Setting it to False says "a missing value means I cannot answer," and on .max() you get NaN back. On .idxmax() there is no NaN label to return, so it raises instead.
numeric_only=False is the default on a data frame, which means Pandas will happily take the maximum of your text columns too. That is the source of the first mistake below.
A worked example, on real data
Here is the Global Coal Plant Tracker, the workbook behind Bamboo Weekly #64. One row per generating unit, 13,906 of them:
import pandas as pd
url = ('https://www.bambooweekly.com/content/files/wp-content/uploads/2024/02/'
'global-coal-plant-tracker-january-2024.xlsx')
units = pd.read_excel(url, sheet_name='Units',
usecols=['Country', 'Region', 'Plant name',
'Status', 'Capacity (MW)', 'Start year'])
operating = units.loc[lambda df_: df_['Status'] == 'operating']
What is the largest coal-burning unit running anywhere in the world?
operating['Capacity (MW)'].max()
1350.0
A number, with no idea whose it is. That is what .idxmax() is for:
operating['Capacity (MW)'].idxmax()
2575
A label — here an integer, because the frame came in with a default index, but a label all the same. Feed it to .loc[] and the answer arrives in full:
operating.loc[operating['Capacity (MW)'].idxmax()]
Country China
Plant name Huaibei Pingshan power station
Capacity (MW) 1350.0
Status operating
Start year 2022.0
Region Asia
Name: 2575, dtype: object
Now let me aggregate to the country level, which gives me a frame whose index is made of names rather than numbers:
by_country = (
operating
.groupby('Country')['Capacity (MW)']
.agg(total='sum', largest_unit='max', units='count')
)
by_country.head()
total largest_unit units
Country
Argentina 495.0 375.0 2
Australia 22403.0 750.0 53
Bangladesh 4775.0 660.0 10
Bosnia and Herzegovina 2090.0 300.0 10
Botswana 732.0 150.0 8
.idxmax() on a data frame runs down every column and returns one label each:
by_country.idxmax()
total China
largest_unit China
units China
dtype: str
China three times over, which surprises nobody. The other end is more interesting — .idxmin() finds the country with the smallest operating fleet, and .loc[] shows it whole:
by_country.loc[by_country['total'].idxmin()]
total 64.0
largest_unit 32.0
units 2.0
Name: Guadeloupe, dtype: float64
Turn the frame sideways and axis='columns' starts earning its keep. Pivoting capacity by status gives one column per status, so asking which column is largest asks which phase of life each country's coal fleet is mostly in:
capacity = (
units
.loc[lambda df_: df_['Status'].isin(['operating', 'construction', 'retired'])]
.pivot_table(index='Country', columns='Status',
values='Capacity (MW)', aggfunc='sum')
)
capacity.idxmax(axis='columns').head(8)
Country
Argentina operating
Australia operating
Austria retired
Bangladesh operating
Belgium retired
Bosnia and Herzegovina operating
Botswana operating
Brazil operating
dtype: str
Austria and Belgium have retired more coal capacity than they still run. The same table read the default way, down the columns, names the champion of each status:
capacity.idxmax()
Status
construction China
operating China
retired United States
dtype: str
Finally, the pairing that matters most in daily work. groupby(...).max() gives you a table of numbers:
operating.groupby('Region')['Capacity (MW)'].max()
Region
Africa 800.0
Americas 1300.0
Asia 1350.0
Europe 1100.0
Oceania 750.0
Name: Capacity (MW), dtype: float64
Whereas groupby(...).idxmax() gives you a table of labels — one row label per group — which you then hand straight back to .loc[] to recover the winning rows themselves:
operating.loc[operating.groupby('Region')['Capacity (MW)'].idxmax(),
['Region', 'Country', 'Plant name', 'Capacity (MW)']]
Region Country Plant name Capacity (MW)
11523 Africa South Africa Kusile power station 800.0
12436 Americas United States Amos Plant 1300.0
2575 Asia China Huaibei Pingshan power station 1350.0
6922 Europe Germany Datteln power station 1100.0
63 Oceania Australia Kogan Creek power station 750.0
That is the biggest operating coal unit on each continent, named, in three lines. groupby().max() could never have told me the plant names, because by the time it has taken the maximum, the rows are gone.
Two mistakes people make
Reaching for .idxmax() when you wanted .max(), or the reverse. One returns the value, the other the label it sits at, and df.loc[…idxmax()] returns the whole row. That distinction, the ties that go silently to the first occurrence, and the all-NaN column that raises on idxmax while max returns NaN, are all on the idxmax and idxmin pages.
.max() on text compares lexically, and does it silently. With numeric_only=False as the default, a data frame of mixed types answers anyway:
units.max()
Country Zimbabwe
Plant name Štavalj Power Station
Capacity (MW) 6300.0
Status shelved
Start year 2037.0
Region Oceania
dtype: object
Zimbabwe is not the largest country and "shelved" is not the largest status — those are just the strings that sort last, and "Štavalj" only wins because Š sorts after Z in Unicode. Nothing warns you. Pass numeric_only=True when you mean numbers.
Where it shows up in Bamboo Weekly
#9: US house prices finds the peak of the index with df['price'].max(), then asks which row it was — the question that leads straight to idxmax, and the reason the two methods are usually typed within a line of each other.
#34: House of Representatives uses groupby('state')['district'].max() to find how many districts each state has — a plain grouped maximum, where the number is the whole answer and no label is needed.
#54: Household debt averages six categories of delinquent loans by year and then compares them, which is the case for axis='columns': the reducer runs across a row rather than down a column.
Practice it
Work through a max exercise, with instant feedback and no signup required: practice.lernerpython.com/bamboo-weekly/max/
Go deeper
If you want the top few rather than the top one, that is nlargest and its twin nsmallest, which beat sorting the whole frame with sort_values. If you want several of these answers at once, agg takes them by name — .agg(['mean', 'min', 'idxmin', 'max', 'idxmax']) gives you the extremes and their locations in a single pass. And since every idxmax ends in a lookup, loc is the method to be fluent in before this one.
More Pandas videos on Python and Pandas with Reuven Lerner.
Bamboo Weekly is the practice. If you want the structured version — full courses with downloadable Jupyter notebooks, plus live Pandas office hours when you get stuck — that is LernerPython+Data. A paid Bamboo Weekly subscription is included with it.
Related methods
.nlargest()— when you want the top few rather than the single largest.idxmax()— when you want to know where the largest value is, not what it is.min()— for the bottom of the range rather than the top
See it on real data
Below are the 23 Bamboo Weekly exercises that use max on real-world data — try each one, then study the worked solution.
- Bamboo Weekly #188: Hurricane season
- Bamboo Weekly #183: Hiring
- Bamboo Weekly #177: European Summer
- Bamboo Weekly #165: Artemis II
- Bamboo Weekly #161: Missiles in Israel
- Bamboo Weekly #157: Government corruption
- Bamboo Weekly #145: Economic indicators
- Bamboo Weekly #142: Hurricanes
- Bamboo Weekly #111: State taxes
- Bamboo Weekly #102: WordPress
- Bamboo Weekly #90: Voter participation
- Bamboo Weekly #85: PACs and parties
- Bamboo Weekly #84: Central banks
- Bamboo Weekly #65: Microplastics
- Bamboo Weekly #56: Rent increases
- Bamboo Weekly #39: WeWork
- Bamboo Weekly #34: House of Representatives
- Bamboo Weekly #29: Auto accidents
- Bamboo Weekly #26: Hot weather
- Bamboo Weekly #19: Working women
- Bamboo Weekly #16: Consumer oil prices
- Bamboo Weekly #9: US house prices
- Bamboo Weekly #7: Bank failures
Part of the Pandas Methods Index. See also practice by skill.