镜像站点 · 本页由第三方 GitHub 只读镜像提供,非 GitHub 官方站点,不接受任何登录或凭据输入。前往 github.com
Skip to content

fix(spark): floor instead of truncate in unix_millis/unix_seconds for pre-epoch timestamp - #26129

Open
harsh-ande wants to merge 1 commit into
apache:mainfrom
harsh-ande:fix-spark-unix-floor
Open

harsh-ande wants to merge 1 commit into
apache:mainfrom
harsh-ande:fix-spark-unix-floor

Conversation

@harsh-ande

Copy link
Copy Markdown

Which issue does this PR close?

Rationale for this change

The Spark-compatible unix_millis and unix_seconds functions return different results from Spark for timestamps before 1970-01-01 that have a sub-unit component. Spark floors the result, but DataFusion truncates toward zero. For example, unix_millis of a timestamp 1 µs before the epoch is -1 in Spark and 0 in DataFusion. Results for non-negative timestamps and exact multiples already match.

The cause is that SparkUnixTimestamp::simplify rewrites the call as a cast to a coarser Timestamp unit, and Arrow's cast truncates toward zero. Spark computes these values with Math.floorDiv (TimestampToLongBase).

What changes are included in this PR?

In SparkUnixTimestamp::simplify (datafusion/spark/src/function/datetime/unix.rs):

  • When the input timestamp unit is finer than the target unit, the raw Int64 ticks are divided by the unit ratio (1e3, 1e6 or 1e9). Then 1 is subtracted when the remainder is negative: ticks / d - CAST(ticks % d < 0 AS BIGINT). This is an exact floor division over the full i64 range, with no intermediate unit conversion, so it adds no new overflow cases.
  • When the input unit is the same as or coarser than the target, the existing direct cast is kept, so results for those inputs don't change.
  • NULL handling doesn't change; all operators involved propagate NULL.

What is the testing strategy for this PR?

New sqllogictest cases in datafusion/sqllogictest/test_files/spark/datetime/unix.slt:

  • Runtime path: a table column (not just literals, so constant folding isn't the only path tested) holding pre-epoch values with and without a remainder, a positive value and NULL. Each is checked for both unix_millis and unix_seconds.
  • Constant-folded path: the same cases written as literals.
  • Other input units: Timestamp(Nanosecond) to millis and micros, Timestamp(Millisecond) to seconds, and an unchanged conversion to a finer unit (Millisecond to micros).
  • Extremes: i64::MIN microseconds to seconds, and a large millisecond value to seconds, to guard against overflow.

Expected values for microsecond inputs came from Spark 4.0.0 (pyspark with spark.sql.session.timeZone=UTC), not from this implementation. Spark has no nanosecond timestamps, so the nanosecond cases use the same floor rule.

Local runs: all spark/* sqllogictests, cargo test -p datafusion-spark, cargo fmt --all -- --check and cargo clippy -p datafusion-spark --all-targets -- -D warnings pass

Are there any user-facing changes?

Yes, as a bug fix. unix_millis and unix_seconds now return the Spark-compatible (floored) value for pre-epoch timestamps with a sub-unit component. No public API changes, so no api change label is needed.


AI assistance: I used an AI coding assistant to help find this divergence (by differential testing against Spark) and to draft the change. A separate AI review pass caught overflow edge cases, which the tests now cover. I reproduced the behavior against Spark 4.0.0 myself and reviewed the change line by line.

@github-actions github-actions Bot added sqllogictest SQL Logic Tests (.slt) spark labels Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

spark sqllogictest SQL Logic Tests (.slt)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[datafusion-spark] unix_millis / unix_seconds truncate toward zero instead of flooring for pre-epoch timestamps

1 participant