Spark DataFrame groupBy and sort in the descending order (pyspark)

python apache-spark dataframe pyspark apache-spark-sql

In PySpark 1.3 sort method doesn't take ascending parameter. You can use desc method instead:

from pyspark.sql.functions import col(group_by_dataframe    .count()    .filter("`count` >= 10")    .sort(col("count").desc()))

or desc function:

from pyspark.sql.functions import desc(group_by_dataframe    .count()    .filter("`count` >= 10")    .sort(desc("count"))

Both methods can be used with with Spark >= 1.3 (including Spark 2.x).

python apache-spark dataframe pyspark apache-spark-sql

Use orderBy:

df.orderBy('column_name', ascending=False)

Complete answer:

group_by_dataframe.count().filter("`count` >= 10").orderBy('count', ascending=False)

python apache-spark dataframe pyspark apache-spark-sql

By far the most convenient way is using this:

df.orderBy(df.column_name.desc())

Doesn't require special imports.

CodeHunter