Convert PySpark dataframe column from list to string

python apache-spark pyspark apache-spark-sql pyspark-sql

While you can use a UserDefinedFunction it is very inefficient. Instead it is better to use concat_ws function:

from pyspark.sql.functions import concat_wsdf.withColumn("test_123", concat_ws(",", "test_123")).show()

+----+----------------+|uuid|        test_123|+----+----------------+|   1|test,test2,test3||   2|test4,test,test6||   3|test6,test9,t55o|+----+----------------+

python apache-spark pyspark apache-spark-sql pyspark-sql

You can create a udf that joins array/list and then apply it to the test column:

from pyspark.sql.functions import udf, coljoin_udf = udf(lambda x: ",".join(x))df.withColumn("test_123", join_udf(col("test_123"))).show()+----+----------------+|uuid|        test_123|+----+----------------+|   1|test,test2,test3||   2|test4,test,test6||   3|test6,test9,t55o|+----+----------------+

The initial data frame is created from:

from pyspark.sql.types import StructType, StructFieldschema = StructType([StructField("uuid",IntegerType(),True),StructField("test_123",ArrayType(StringType(),True),True)])rdd = sc.parallelize([[1, ["test","test2","test3"]], [2, ["test4","test","test6"]],[3,["test6","test9","t55o"]]])df = spark.createDataFrame(rdd, schema)df.show()+----+--------------------+|uuid|            test_123|+----+--------------------+|   1|[test, test2, test3]||   2|[test4, test, test6]||   3|[test6, test9, t55o]|+----+--------------------+

python apache-spark pyspark apache-spark-sql pyspark-sql

As of version 2.4.0, you can use array_join.Spark docs

from pyspark.sql.functions import array_joindf.withColumn("test_123", array_join("test_123", ",")).show()

CodeHunter

Convert PySpark dataframe column from list to string

Recent Posts

How can I color dots in a xy scatterplot according to column value?

How to update a claim in ASP.NET Identity?

What does {0} mean when initializing an object?

Accessing members of items in a JSONArray with Java

How to log SQL statements in Spring Boot?

Powershell Get-WebSite name parameter is ignored

How to detect scroll to bottom of html element

Java synchronized method

How to test controllers with CodeIgniter?

Detect Visual Composer

Matplotlib: Specify format of floats for tick labels

Rails join a list of strings with commas and "and" before the last