max returns the largest value. idxmax returns the index of that largest value.
Which row was that? You called .max(), you got a number back, and now you are stuck — because .max() answers "how big," and throws away everything else on the way. .idxmax() answers the other half of the question, returning the index label of the row where that maximum lives.
The label is the whole point, because a label is exactly what .loc[] takes. Put the two together and you get the line most people are really here for:
df.loc[df['column'].idxmax()]
That is "give me the entire row where this column is highest," without sorting the frame and without a boolean mask comparing a column against its own maximum. On a series .idxmax() returns one label; on a data frame, one label per column.
Official documentation: Series.idxmax and DataFrame.idxmax.
The arguments that earn their keep
df.idxmax(axis='index', # down each column (default), or across each row
skipna=True, # ignore NaN, or refuse to answer
numeric_only=False) # every column, or only the numeric ones
axis changes the question, not just the answer. The default runs down each column and gives you a row label per column; axis='columns' runs across each row and gives you a column name per row. Which year was hottest, versus which month was hottest.
A worked example, on real data
NASA's GISTEMP record: one row per year, one column per month, plus J-D for the annual mean, all in degrees Celsius above the 1951–1980 baseline. GISS rebuilds the file every month, so your numbers will run a little past mine.
import pandas as pd
url = 'https://data.giss.nasa.gov/gistemp/tabledata_v4/GLB.Ts+dSST.csv'
temps = pd.read_csv(url, skiprows=1, index_col='Year', na_values='***',
usecols=['Year', 'Jan', 'Feb', 'Mar', 'Apr', 'May', 'Jun',
'Jul', 'Aug', 'Sep', 'Oct', 'Nov', 'Dec', 'J-D'])
temps['J-D'].max()
1.28
An anomaly with no address. .idxmax() supplies one:
temps['J-D'].idxmax()
2024
And because that is a label, .loc[] will take it and give the year back whole:
temps.loc[temps['J-D'].idxmax()]
Jan 1.25
Feb 1.44
Mar 1.39
Apr 1.31
May 1.16
Jun 1.23
Jul 1.20
Aug 1.30
Sep 1.21
Oct 1.35
Nov 1.30
Dec 1.27
J-D 1.28
Name: 2024, dtype: float64
Read the same table sideways and the answers become month names:
months = temps.drop(columns='J-D')
months.idxmax(axis='columns').tail(6)
Year
2021 Oct
2022 Mar
2023 Sep
2024 Feb
2025 Jan
2026 Mar
dtype: str
Which is an ordinary series of column names, so it counts like anything else:
months.idxmax(axis='columns').value_counts().head(4)
Jan 21
Feb 21
Dec 20
Nov 19
Name: count, dtype: int64
Against a mid-century baseline it is the winter months that run hottest, which says how fast northern winters are changing rather than anything about absolute temperature.
The trick worth knowing: first True
.idxmax() on a boolean series finds the first row where a condition holds. True is 1, False is 0, and ties go to the first occurrence, so the maximum of a boolean series is the earliest True:
(temps['J-D'] > 1.0).idxmax()
2016
2016 was the first year the global anomaly cleared a full degree. No sorting, no .cumsum(), no loop.
The trap arrives with it. If the condition is never satisfied, every value ties at False, and you get the first label in the series:
(temps['J-D'] > 5.0).idxmax()
1880
Nothing warns you that 1880 is a non-answer. Check .any() first whenever a match is not guaranteed.
Three mistakes people make
A label is not a position. .idxmax() said 2024; .argmax(), the positional twin, says 144. Hand 2024 to .iloc[] and you get IndexError: single positional indexer is out-of-bounds, which is the lucky outcome. The dangerous case is an index of integers that could pass for positions — district numbers, ZIP codes — where .iloc[] quietly returns some other row. Pair .idxmax() with .loc[], and .argmax() with .iloc[].
Ties are broken silently, always for the first. The count above puts January on top with 21 years — and February also has 21. Ask which month wins and Pandas answers without a flicker:
months.idxmax(axis='columns').value_counts().idxmax()
'Jan'
When a tie is plausible, compare against .max() with a boolean mask and count how many rows come back.
All-NaN raises, where .max() shrugs. The current year is only partly reported, so its later months are missing. .max() returns nan; .idxmax() has no label to point at, and raises ValueError: Encountered all NA values. skipna=False on a column with even one missing value raises too. Worst of all, inside a groupby one empty group takes down the whole call: ValueError: idxmax with skipna=True encountered all NA values in a group.
Where it shows up in Bamboo Weekly
#74: UK elections pivots constituencies into parties by region, then calls .idxmax() so each column reports its winning party rather than its winning count. Labour in every region but one.
#48: Aviation accidents builds a 17,489-row table of operators by year with value_counts().unstack(), and .idxmax() names the operator with the most incidents in each of the 27 years.
#27: Young voters is the series case, on one row pulled out of a MultiIndex frame: .loc['Definitely will be voting', 'Race/Ethnicity Category'].idxmax() returns the demographic label directly.
#39: WeWork has a DatetimeIndex, so we_df['Close'].idxmax() comes back as Timestamp('2021-10-25 00:00:00') — the return value is whatever your index labels happen to be.
Practice it
Work through an idxmax exercise, with instant feedback and no signup required: practice.lernerpython.com/bamboo-weekly/idxmax/
Go deeper
max, min, idxmax and idxmin covers the family together, and idxmin is this method pointed downward, with idioms of its own. For the top few rather than the top one, reach for nlargest rather than sort_values. And since every idxmax ends in a lookup, loc is the method to be fluent in first.
More Pandas videos on Python and Pandas with Reuven Lerner.
Bamboo Weekly is the practice. If you want the structured version — full courses with downloadable Jupyter notebooks, plus live Pandas office hours when you get stuck — that is LernerPython+Data. A paid Bamboo Weekly subscription is included with it.
Related methods
.max()— when you want the value itself rather than the label it sits at.nlargest()— when you want the top few rows rather than the single winning label.idxmin()— for the other end of the range
See it on real data
Below are the 15 Bamboo Weekly exercises that use idxmax on real-world data — try each one, then study the worked solution.
- Bamboo Weekly #177: European Summer
- Bamboo Weekly #119: Python conferences
- Bamboo Weekly #74: UK elections
- Bamboo Weekly #54: Household debt
- Bamboo Weekly #48: Aviation accidents
- Bamboo Weekly #39: WeWork
- Bamboo Weekly #34: House of Representatives
- Bamboo Weekly #31: Poverty
- Bamboo Weekly #30: Uncertainty
- Bamboo Weekly #29: Auto accidents
- Bamboo Weekly #27: Young voters
- Bamboo Weekly #19: Working women
- Bamboo Weekly #10: Oil prices
- Bamboo Weekly #9: US house prices
- Bamboo Weekly #1: Government corruption
Part of the Pandas Methods Index. See also practice by skill.