Skip to content

pandas idxmax

max returns the largest value. idxmax returns the index of that largest value.

Which row was that? You called .max(), you got a number back, and now you are stuck — because .max() answers "how big," and throws away everything else on the way. .idxmax() answers the other half of the question, returning the index label of the row where that maximum lives.

The label is the whole point, because a label is exactly what .loc[] takes. Put the two together and you get the line most people are really here for:

df.loc[df['column'].idxmax()]

That is "give me the entire row where this column is highest," without sorting the frame and without a boolean mask comparing a column against its own maximum. On a series .idxmax() returns one label; on a data frame, one label per column.

Official documentation: Series.idxmax and DataFrame.idxmax.

The arguments that earn their keep

df.idxmax(axis='index',        # down each column (default), or across each row
          skipna=True,         # ignore NaN, or refuse to answer
          numeric_only=False)  # every column, or only the numeric ones

axis changes the question, not just the answer. The default runs down each column and gives you a row label per column; axis='columns' runs across each row and gives you a column name per row. Which year was hottest, versus which month was hottest.

A worked example, on real data

NASA's GISTEMP record: one row per year, one column per month, plus J-D for the annual mean, all in degrees Celsius above the 1951–1980 baseline. GISS rebuilds the file every month, so your numbers will run a little past mine.

import pandas as pd

url = 'https://data.giss.nasa.gov/gistemp/tabledata_v4/GLB.Ts+dSST.csv'

temps = pd.read_csv(url, skiprows=1, index_col='Year', na_values='***',
                    usecols=['Year', 'Jan', 'Feb', 'Mar', 'Apr', 'May', 'Jun',
                             'Jul', 'Aug', 'Sep', 'Oct', 'Nov', 'Dec', 'J-D'])

temps['J-D'].max()
1.28

An anomaly with no address. .idxmax() supplies one:

temps['J-D'].idxmax()
2024

And because that is a label, .loc[] will take it and give the year back whole:

temps.loc[temps['J-D'].idxmax()]
Jan    1.25
Feb    1.44
Mar    1.39
Apr    1.31
May    1.16
Jun    1.23
Jul    1.20
Aug    1.30
Sep    1.21
Oct    1.35
Nov    1.30
Dec    1.27
J-D    1.28
Name: 2024, dtype: float64

Read the same table sideways and the answers become month names:

months = temps.drop(columns='J-D')

months.idxmax(axis='columns').tail(6)
Year
2021    Oct
2022    Mar
2023    Sep
2024    Feb
2025    Jan
2026    Mar
dtype: str

Which is an ordinary series of column names, so it counts like anything else:

months.idxmax(axis='columns').value_counts().head(4)
Jan    21
Feb    21
Dec    20
Nov    19
Name: count, dtype: int64

Against a mid-century baseline it is the winter months that run hottest, which says how fast northern winters are changing rather than anything about absolute temperature.

The trick worth knowing: first True

.idxmax() on a boolean series finds the first row where a condition holds. True is 1, False is 0, and ties go to the first occurrence, so the maximum of a boolean series is the earliest True:

(temps['J-D'] > 1.0).idxmax()
2016

2016 was the first year the global anomaly cleared a full degree. No sorting, no .cumsum(), no loop.

The trap arrives with it. If the condition is never satisfied, every value ties at False, and you get the first label in the series:

(temps['J-D'] > 5.0).idxmax()
1880

Nothing warns you that 1880 is a non-answer. Check .any() first whenever a match is not guaranteed.

Three mistakes people make

A label is not a position. .idxmax() said 2024; .argmax(), the positional twin, says 144. Hand 2024 to .iloc[] and you get IndexError: single positional indexer is out-of-bounds, which is the lucky outcome. The dangerous case is an index of integers that could pass for positions — district numbers, ZIP codes — where .iloc[] quietly returns some other row. Pair .idxmax() with .loc[], and .argmax() with .iloc[].

Ties are broken silently, always for the first. The count above puts January on top with 21 years — and February also has 21. Ask which month wins and Pandas answers without a flicker:

months.idxmax(axis='columns').value_counts().idxmax()
'Jan'

When a tie is plausible, compare against .max() with a boolean mask and count how many rows come back.

All-NaN raises, where .max() shrugs. The current year is only partly reported, so its later months are missing. .max() returns nan; .idxmax() has no label to point at, and raises ValueError: Encountered all NA values. skipna=False on a column with even one missing value raises too. Worst of all, inside a groupby one empty group takes down the whole call: ValueError: idxmax with skipna=True encountered all NA values in a group.

Where it shows up in Bamboo Weekly

#74: UK elections pivots constituencies into parties by region, then calls .idxmax() so each column reports its winning party rather than its winning count. Labour in every region but one.

#48: Aviation accidents builds a 17,489-row table of operators by year with value_counts().unstack(), and .idxmax() names the operator with the most incidents in each of the 27 years.

#27: Young voters is the series case, on one row pulled out of a MultiIndex frame: .loc['Definitely will be voting', 'Race/Ethnicity Category'].idxmax() returns the demographic label directly.

#39: WeWork has a DatetimeIndex, so we_df['Close'].idxmax() comes back as Timestamp('2021-10-25 00:00:00') — the return value is whatever your index labels happen to be.

Practice it

Work through an idxmax exercise, with instant feedback and no signup required: practice.lernerpython.com/bamboo-weekly/idxmax/

Go deeper

max, min, idxmax and idxmin covers the family together, and idxmin is this method pointed downward, with idioms of its own. For the top few rather than the top one, reach for nlargest rather than sort_values. And since every idxmax ends in a lookup, loc is the method to be fluent in first.

More Pandas videos on Python and Pandas with Reuven Lerner.

Bamboo Weekly is the practice. If you want the structured version — full courses with downloadable Jupyter notebooks, plus live Pandas office hours when you get stuck — that is LernerPython+Data. A paid Bamboo Weekly subscription is included with it.

See it on real data

Below are the 15 Bamboo Weekly exercises that use idxmax on real-world data — try each one, then study the worked solution.

Part of the Pandas Methods Index. See also practice by skill.