CertSafari
    Databricks Certified Associate Developer for Apache Spark· Lessons

    Domain 3 · Lesson 15/32

    Unix Epoch to Date String in PySpark: from_unixtime, unix_timestamp, to_date

    Manipulate and utilize Date data type, such as Unix epoch to date string, and extract date component.

    10 min read
    3.12% of exam
    7 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Convert a column of Unix epoch seconds into a formatted timestamp string with from_unixtime
    • Turn a date or timestamp string back into epoch seconds with unix_timestamp, and predict when it returns null
    • Produce a true DateType column with to_date, and render any date as a custom string with date_format
    • Read and write datetime patterns, telling MM from mm and d from D

    Key concept

    Unix epoch seconds — A Unix time is a count of seconds since 1970-01-01 00:00:00 UTC. Spark's epoch functions convert between that number, a formatted string, and the DATE or TIMESTAMP types, and the type each one returns is what decides your next step.

    1.Three representations of the same moment

    Event data often arrives with time stored as an integer, such as 1428476400. That integer is a Unix time: seconds counted from 1970-01-01 00:00:00 UTC. People don't read it easily, and you can't group it by calendar day without converting it first. In Spark a single moment can be held three ways: as a number of epoch seconds (a long), as a string like 2015-04-08 12:12:12, or as a typed value, either DATE or TIMESTAMP.

    The DATE type holds only a calendar day. It stores year, month and day and has no time zone, and the supported range runs from June 23 -5877641 CE to July 11 +5881580 CE. In SQL you write a DATE literal as DATE'2020-12-31'. If the literal isn't a valid date, Databricks raises an error.

    Checkpoint 1 of 7· Check yourself

    Which fields does a Spark DATE value hold?

    Most mistakes on this objective come from getting the return type wrong. The table below lists the conversion functions in the PySpark functions reference that move between epoch numbers, strings and typed values.

    Epoch and date conversion functions and what each one returns
    FunctionInputOutput
    from_unixtime(timestamp[, format])Seconds since the Unix epochFormatted string
    unix_timestamp([timestamp, format])Time string with a patternUnix time in seconds
    timestamp_seconds(col)Seconds since the Unix epochTimestamp
    date_from_unix_date(days)Days since 1970-01-01Date
    unix_date(col)DateDays since 1970-01-01
    to_date(col[, format])Column such as a stringDateType
    date_format(date, format)Date, timestamp or stringString in the given format

    Sources12

    2.from_unixtime: epoch seconds to a timestamp string

    from_unixtime(timestamp, format) takes a column of Unix time values and returns a string. The format argument is an optional literal string that defaults to yyyy-MM-dd HH:mm:ss. Pass a different pattern, such as yyyy-MM-dd or dd.MM.yyyy, to get a different layout.

    The string shows the moment in the session's time zone, not in UTC. The documented example therefore sets spark.sql.session.timeZone to America/Los_Angeles before it runs and unsets it afterwards. Run the same epoch value under two session time zones and you can get two different strings.

    Converting a unix_time column with the default formatpython
    from pyspark.sql import functions as dbf
    df = spark.createDataFrame([(1428476400,)], ['unix_time'])
    df.select('*', dbf.from_unixtime('unix_time')).show()

    Checkpoint 2 of 7· Fill the gap

    Which function completes this sample so that the epoch-seconds column becomes a formatted timestamp string?

    from pyspark.sql import functions as dbf
    df = spark.createDataFrame([(1428476400,)], ['unix_time'])
    df.select('*', dbf. ? ('unix_time')).show()

    Checkpoint 3 of 7· Exam question

    A pipeline stores event times as Unix epoch seconds in a `LongType` column named `event_epoch`. An analyst needs a new column `event_date` holding the date as a string in `yyyy-MM-dd` format. Which code correctly creates it? ```python df.withColumn("event_date", ___) ```

    Sources3

    3.unix_timestamp: strings back to epoch seconds

    unix_timestamp(timestamp, format) does the reverse conversion. It parses a time string with the given pattern, yyyy-MM-dd HH:mm:ss by default, and returns Unix time in seconds as a long integer. Parsing uses the default time zone and locale. A string that doesn't parse produces null, not an error, so a wrong pattern can turn a whole column into nulls without any warning. With no argument, unix_timestamp() returns the current timestamp.

    The documentation gives two examples. Under the default pattern, the string 2015-04-08 12:12:12 parses to 1428520332. With the pattern yyyy-MM-dd, the date-only string 2015-04-08 parses to 1428476400, the same number used in the from_unixtime example.

    Parsing a date-only string by supplying a matching patternpython
    import pyspark.sql.functions as sf
    df = spark.createDataFrame([('2015-04-08',)], ['dt'])
    df.select('*', sf.unix_timestamp('dt', 'yyyy-MM-dd')).show()

    Checkpoint 4 of 7· Check yourself

    A column holds strings like '08/04/2015'. You call unix_timestamp on it without a format argument. What happens?

    Sources4

    4.to_date for a real DATE, date_format for display

    Two functions finish an epoch-to-date conversion. Which one you need depends on whether the result should be a typed value or text.

    to_date(col, format) returns a column of DateType. If you omit the format, it follows the normal casting rules and is equivalent to col.cast("date"). If you supply a format, it parses with that pattern. Both calls in the example below turn the timestamp string 1997-02-28 10:30:00 into a date. A string from from_unixtime can be passed to to_date in the same way to get a DATE column.

    date_format(date, format) goes the other direction. It accepts a date, timestamp or string and returns a string in the pattern you give, for example MM/dd/yyyy or dd.MM.yyyy, which produces strings like 18.03.1993. Use it to produce output text, such as a report label or a file-name fragment.

    to_date with and without an explicit patternpython
    from pyspark.sql import functions as dbf
    df = spark.createDataFrame([('1997-02-28 10:30:00',)], ['ts'])
    df.select('*', dbf.to_date(df.ts)).show()
    df.select('*', dbf.to_date('ts', 'yyyy-MM-dd HH:mm:ss')).show()
    date_format rendering a date as MM/dd/yyyy textpython
    from pyspark.sql import functions as dbf
    df = spark.createDataFrame([('2015-04-08',), ('2024-10-31',)], ['dt'])
    df.select("*", dbf.typeof('dt'), dbf.date_format('dt', 'MM/dd/yyyy')).show()

    Checkpoint 5 of 7· Match them up

    Match each function to the type it returns

    Tap a term, then the definition that fits it.

    Sources56

    5.Datetime pattern letters

    unix_timestamp, date_format, from_unixtime and to_date all read the same pattern letters. Letter case matters. The table lists the letters that most often get mixed up.

    Pattern letters that differ only by case
    SymbolMeaningExample
    Mmonth-of-year7; 07; Jul; July
    mminute-of-hour30
    dday-of-month28
    Dday-of-year189
    Hhour-of-day (0-23)0
    hclock-hour-of-am-pm (1-12)12
    Eday-of-weekTue; Tuesday

    The number of times a letter repeats also changes the output. For text fields such as E, fewer than four letters give the short form (Mon), exactly four give the full form (Monday), and five or more fail. Month follows the number/text rule: M prints 1 through 9 without padding, MM zero-pads, MMM gives Jan and MMMM gives January. For years, yy prints the last two digits. When parsing, yy uses a base value of 2000, so the year it produces always falls between 2000 and 2099.

    Checkpoint 6 of 7· Check yourself

    A colleague writes date_format(col, 'yyyy-mm-dd') and the middle field shows values like 30 or 00 instead of the month. Why?

    Checkpoint 7 of 7· Exam question

    A raw log column `log_time` (string) stores values like `"2024-03-15 14:30:00"` in the format `yyyy-MM-dd HH:mm:ss`. A downstream system requires a `LongType` column `epoch_seconds` holding the number of seconds since the Unix epoch. Which code correctly creates it? ```python df.withColumn("epoch_seconds", ___) ```

    Sources7

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.from_unixtime returns a DATE or TIMESTAMP column you can use directly in date arithmetic.Why is that wrong?

      It returns a formatted string. Wrap it in to_date, or use timestamp_seconds, if you need a typed value.

      Covered in from_unixtime: epoch seconds to a timestamp string

    2. 2.unix_timestamp throws an error when a string doesn't match the pattern.Why is that wrong?

      A failed parse returns null, so a wrong pattern quietly turns the column into nulls.

      Covered in unix_timestamp: strings back to epoch seconds

    3. 3.In a datetime pattern, mm means month.Why is that wrong?

      Lowercase m is minute-of-hour. Month-of-year is uppercase M or MM.

      Covered in Datetime pattern letters

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Represents values comprising values of fields year, month, and day, without a time-zone.”
      ↩︎ Three representations of the same moment
    2. 2.
      “timestamp_seconds(col) | Converts the number of seconds from the Unix epoch (1970-01-01T00:00:00Z) to a timestamp.”
      ↩︎ Three representations of the same moment
      “date_from_unix_date(days) | Create date from the number of days since 1970-01-01.”
      ↩︎ Three representations of the same moment
    3. 3.
      “a string representing the timestamp of that moment in the current system time zone in the given format”
      ↩︎ from_unixtime: epoch seconds to a timestamp string
      “format to use to convert to (default: yyyy-MM-dd HH:mm:ss)”
      ↩︎ from_unixtime: epoch seconds to a timestamp string
      “Converts the number of seconds from unix epoch (1970-01-01 00:00:00 UTC) to a string representing the timestamp of that moment”
      ↩︎ Key concept
      “pyspark.sql.Column: formatted timestamp as string.”
      ↩︎ Exam trap 1
      “pyspark.sql.Column: formatted timestamp as string.”
      ↩︎ Prediction
    4. 4.
      “Convert time string with given pattern ('yyyy-MM-dd HH:mm:ss', by default) to Unix time stamp (in seconds)”
      ↩︎ unix_timestamp: strings back to epoch seconds
      “If timestamp is None, then it returns current timestamp.”
      ↩︎ unix_timestamp: strings back to epoch seconds
      “pyspark.sql.Column: unix time as long integer.”
      ↩︎ unix_timestamp: strings back to epoch seconds
      “using the default timezone and the default locale, returns null if failed.”
      ↩︎ Exam trap 2
      “using the default timezone and the default locale, returns null if failed.”
      ↩︎ Checkpoint
    5. 5.
      “By default, it follows casting rules to pyspark.sql.types.DateType if the format is omitted. Equivalent to col.cast("date").”
      ↩︎ to_date for a real DATE, date_format for display
      “pyspark.sql.Column: date value as pyspark.sql.types.DateType type.”
      ↩︎ Checkpoint
    6. 6.
      “Converts a date/timestamp/string to a value of string in the format specified by the date format given by the second argument.”
      ↩︎ to_date for a real DATE, date_format for display
    7. 7.
      “Exactly 4 pattern letters will use the full text form”
      ↩︎ Datetime pattern letters
      “For parsing, this will parse using the base value of 2000, resulting in a year within the range 2000 to 2099 inclusive.”
      ↩︎ Datetime pattern letters
      “M/L | month-of-year | month | 7; 07; Jul; July”
      ↩︎ Exam trap 3
      “m | minute-of-hour | number(2) | 30”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Extract Date Components in PySpark: year, month, dayofweek, extract

    Spotted a mistake, or was something unclear? Tell us.