Skip to content

Unify random source across stories - #127

Draft
khokhlov962 wants to merge 18 commits into
WebKit:tentative-1-4-testsfrom
khokhlov962:random_seed
Draft

khokhlov962 wants to merge 18 commits into
WebKit:tentative-1-4-testsfrom
khokhlov962:random_seed

Conversation

@khokhlov962

@khokhlov962 khokhlov962 commented Oct 1, 2026 •

Copy link
Copy Markdown

Wrong branch, please ignore

Victor Hugo Vianna Silva and others added 18 commits September 17, 2026 17:57
Change-Id: Ib803f9300a2b2dd83e6e96caf126211ceac9b848
Change-Id: Iafed33715dba9480ac34d4735d87e14131c96915
Trace analysis needs the exact ramp window per subtest. Intervals derived
from iframe navigation overshoot it by 14-20%, and the overshoot cannot be
corrected with a fixed offset because tier sampling is a complexity search
that varies with device speed.

Emit marks at the three transitions that already exist in the control flow
(end of warmup, end of tier sampling, stop), pinned to the controller's own
timestamps, plus measure spans between them. Every call is guarded: an
exception escaping _animateLoop() would stop the requestAnimationFrame chain
and hang the run with no error reported anywhere.

Also report what the controller is doing inside that window, which is
otherwise invisible in a trace: the complexity each ramp interval attempts
along with the frame length it produced, and the change point of each ramp's
regression, which is the complexity at which the target frame rate stops
being met and the value the score is derived from. Both ride along as mark
detail, which Chrome serializes only while blink.user_timing is enabled, so
untraced scoring runs are unaffected.

TAG=agy
CONV=82c228f8-31a9-480b-ba3b-ccc2a5eb2331

Change-Id: Ibdc75af1525748f308a22e57095b0518c207cc0e
https://screenshot-v2.corp.google.com/5lm6dv2klfr20

The 8 tests from the tentative 1.4 suite (Stories, Alice, Chess,
Map Zoomer, Sheets, Départements, Dashboard, Filtering) are already
the primary MotionMark suite in resources/runner/tests.js.
Keeping them as a separate Tentative 1.4 suite in debug-runner/tests.js
duplicates all 8 tests when opening developer.html.

Change-Id: I8c08082f95d510016ded94f1769370cfeb17b9f0
High-level overview:
The current Chess subtest increments `--anim-value` by integer steps
(`++this._animValue`) in `ChessStage.animate()` and feeds it into
`calc(mod(var(--anim-input), 400) * 1grad)` with a static integer
`--random` offset (`0..100`) per leaf, while using a fixed 5-color OKLCH
palette on `#stage`.

Why this is more realistic and less gameable:
Real `requestAnimationFrame`-driven animations advance using elapsed
timestamps (`performance.now()`) rather than frame-counter increments
whose visual speed varies with display refresh rate. Because `--anim-
input`
in the current version is strictly an integer modulo 400, every
`.leaf.type-3` conic gradient and `.leaf.type-4` linear gradient only
ever
visits 400 discrete gradient angles, which would allow a hypothetical
gradient-parameter or shader cache keyed on resolved gradient
stops/angles
to hit after 400 frames. Advancing `--anim-value` from elapsed
`performance.now()` time (as an unrounded floating-point number) makes
the
animation rate frame-rate-independent and increases the gradient angle
state space beyond what is practically cacheable, while modulating
`--palette-color-2..5` on `#stage` prevents static palette memoization
and
keeps `.leaf` `contain: strict` and static glyph masks intact.

Summary of outcomes and changes:
- Update `ChessStage.animate()` to compute `--anim-value` purely from
  elapsed time (`(performance.now() - this._startTime) * 0.06`)
  without a
  frame counter or fixed-decimal truncation.
- Dynamically modulate `--palette-color-2..5` OKLCH lightness and hue
  channels on `#stage` via `sin(calc(var(--anim-value) * ...))`, which
  flows into `.leaf.type-3`/`.leaf.type-4` gradients, `@keyframes fade`,
  and `.leaf.type-1` `color-mix()` styles.
- Keep `contain: strict` and static glyph sizing on `.leaf` elements.

Benchmark impact (Chrome Android, Pixel 10 Pro XL, 60 Hz, N=4 paired
runs):
- Chess score: 17.27 [25.00, 14.56, 14.20, 15.31] -> 16.91 [18.61,
  16.49,
  17.35, 15.19] (-2.1%, no measurable change within run-to-run
  variance).

Change-Id: I25b730e6dea5fcec3968c807487811e9d70961f4
High-level overview:
The current Filtering subtest leaves 2 of the 6 SVG filter cells
(`#neon`
and `#lights`) with completely static filter primitive attributes and
static source graphic styles (`color: red`, `-webkit-text-stroke-color:
blue`, `color: pink`) after frame 0, and declares `--anim-value` with
`syntax: "<integer>"` (stepping `--hue-angle`, `--color1`, and
`--color2`
through 360 integer values on the filtered source graphics of `#zoom`
and
`#signage`).

Why this is more realistic and less gameable:
In active filter animations, all 6 filter cells are intended to perform
per-frame filter rasterization rather than leaving a third of the grid
100% static after frame 0. As shown by the benchmark measurement below,
leaving `#neon` and `#lights` static allowed the pre-change subtest to
score 2.11x (+111%) higher on Chrome than when all 6 cells actively
animate every frame. In addition, changing `@property --anim-value` from
`<integer>` to `<number>` increases the number of distinct `hsl()` input
colors rendered into the `#zoom` and `#signage` source graphics beyond
360
integer steps, reducing the practicality of hypothetical input-keyed
filter output caching while preserving `#watery`'s static `<feTurbulence
seed="20">` so upstream filter-DAG partial-invalidation caches remain
rewarded.

Summary of outcomes and changes:
- Add continuous SMIL `<animate>` elements to `#neon` (`hueRotate`
  from 0
  to 360) and `#lights` (`feDropShadow` `stdDeviation` 4;10;4) so all 6
  filter cells perform per-frame SVG filter rasterization.
- Change `@property --anim-value` from `<integer>` to `<number>` so the
  filtered source graphic `color` inputs of `#zoom` and `#signage` vary
  with floating-point `--hue-angle` steps rather than 360 integers.
- Preserve static `feTurbulence` in `#watery` to reward upstream
  filter-DAG caching.

Benchmark impact (Chrome Android, Pixel 10 Pro XL, 60 Hz, N=2 paired
runs):
- Filtering score: 2.74 [2.77, 2.71] -> 1.30 [1.32, 1.27] (-52.6%), as
  `#neon` and `#lights` are now re-rasterized every frame instead of
  remaining static after frame 0.

Change-Id: Iba7708c3d11a47ad7d6feebf3a3dc2a182b48cf7
High-level overview:
In the current Stories subtest, each `.box` card scales its background
`<img>` (`scale: var(--image-scale)`) inside an `overflow: clip;
border-radius: 16px; opacity: 0.95` container while the foreground
subtree
(`.badge` with `filter: drop-shadow(...)`, `.shadow`, and `.overlay`
text)
has no dependency on `--image-scale` and remains 100% static after
frame 0.
Furthermore, `BoxItem` initializes `new RampAnimator(1, 1.2, 5000,
Stage.random(0, 1))`, passing a random phase offset of only `0..1 ms`
into
a `5000 ms` ramp so every card on screen has nearly the exact same
`--image-scale` value on any given frame.

Why this is more realistic and less gameable:
Because the foreground elements inside `.box` never change after frame
0 in
the current version, an engine could cache the foreground subtree's
rasterization from frame 0 (including the `.badge` `drop-shadow` filter
and text shadow) and reuse that cached texture when compositing the
`opacity: 0.95` card group each frame. In addition, because
`Stage.random(0,
1)` only offsets cards by at most 1 ms out of 5000 ms, all cards in a
frame share virtually identical `--image-scale` values, allowing
cross-card
raster cache sharing within a frame. Fixing the phase offset to
`Stage.random(0, 5000)` gives each card an independent phase across
the 5s
ramp, and animating the `.badge` ring (`conic-gradient`) and `drop-
shadow`
vertical offset via `--image-scale` ensures each card's foreground layer
varies per frame.

Summary of outcomes and changes:
- Fix `new RampAnimator(1, 1.2, 5000, Stage.random(0, 1))` to
  `Stage.random(0, 5000)` in `stories.js` so cards have distinct
  `--image-scale` phases across the 5000 ms ramp.
- Replace the static solid `#084A97` `.badge` border in `stories.html`
  with a rotating multi-color `conic-gradient` border-box ring and
  animate
  its `drop-shadow` vertical offset using `--image-scale`.

Benchmark impact (Chrome Android, Pixel 10 Pro XL, 60 Hz, N=2 paired
runs):
- Stories score: 5.37 [5.23, 5.50] -> 5.58 [5.41, 5.75] (+3.9%, no
  measurable change within run-to-run variance).

Change-Id: I0beb6379c48a656c81cb7daf626f0d4c190fb4fe
High-level overview:
In the current Departements subtest, `#drawWedge()` clips each spoke to
`wedgePath` before filling `areaWedgePath` and `populationWedgePath`
that
have radii `<= endcapRadius` and share the exact same `[startAngle,
endAngle]` radial boundaries (so nothing is ever clipped by
`ctx.clip()`).
In `#drawMap()`, `ctx.clip(wedgePath)` is called before
`ctx.fill(wedgePath)`
even though only the subsequent `spriteSheet.drawCellAtIndex()` call
needs
the clip. In addition, each spoke allocates 6 fresh `Path2D` objects
and a
fresh `CanvasPattern` on every frame.

Why this is more realistic and less gameable:
In production Canvas 2D charts, clip regions actively crop geometry that
extends beyond the clip boundary, and static spoke/arrowhead paths and
patterns are reused across frames via CTM transforms (`ctx.rotate()`).
In
the current harness, because `areaWedgePath` and `populationWedgePath`
in
`#drawWedge()` are geometrically contained within `wedgePath`, a
hypothetical `CanvasRenderingContext2D` fast path that checks whether
the
draw path is contained in the active clip path could elide
`ctx.clip()` in
`#drawWedge()` completely. Padding the angular span of `areaWedgePath`
and
`populationWedgePath` proportionally (`+-0.15 * this.wedgeAngleRadians`)
makes `ctx.clip()` perform genuine subtractive clipping along the radial
spoke edges at constant overdraw across all spoke complexities, with
near-identical visual output (differing only in clipped vs. filled
radial
edge antialiasing).

Summary of outcomes and changes:
- Pad the angular span (`+-0.15 * this.wedgeAngleRadians`) of
  `areaWedgePath` and `populationWedgePath` in `#drawWedge()` so
  `ctx.clip(this._clipWedgePath)` actively clips the radial edges at a
  constant overdraw ratio across all spoke counts.
- Move `ctx.fill(this._mapWedgePath)` before
  `ctx.clip(this._mapWedgePath)`
  in `#drawMap()` so the background fill is not self-clipped while the
  following `spriteSheet.drawCellAtIndex()` remains clipped to the
  wedge.
- Pre-allocate canonical origin-centered `Path2D` objects
  (`_clipWedgePath`, `_mapWedgePath`, `_arrowHeadPath`) rotated via CTM
  and remove the unused `#pathForWedge()` helper, cutting per-spoke
  `Path2D` allocations from 6 to 3 per frame.
- Cache `this._cachedPattern ??= ctx.createPattern(...)` across frames
  instead of recreating the `CanvasPattern` per spoke per frame.

Benchmark impact (Chrome Android, Pixel 10 Pro XL, 60 Hz, N=2 paired
runs):
- Departements score: 34.18 [34.21, 34.15] -> 33.51 [32.93, 34.08]
  (-2.0%,
  no measurable change within run-to-run variance).

Change-Id: I268ba81cd81a58cd30826dc6df7d37ab23da176c
In Regression._windowedFit(), bestComplexity is initialized to 0 and
only updated when a sliding window satisfies
error >= kAllowedErrorFactor (0.9). If a device fails to sustain >= 90%
of the target frame rate even at the lowest complexity in the regression
sample set (minComplexity, after findRegression()'s 2.5% tail trim),
bestComplexity remains 0, which causes the subtest score and the suite
geometric mean score to collapse to 0.

Allow windows at minComplexity (or the first full window via
bestComplexity == 0 when fewer than windowSize samples sit at
minComplexity) to update bestComplexity with
adjustedComplexity = averageComplexity * Math.min(1.0, error) regardless
of kAllowedErrorFactor, while still enforcing
error >= kAllowedErrorFactor for windows above minComplexity and
preserving the strict break behavior.

Fallback semantics and visible effects:
- In 'window' mode, when no window reaches kAllowedErrorFactor, this
  returns the max adjustedComplexity across windows at minComplexity
  (staying continuous across the 0.9 threshold and matching the
  max-over-windows rule used above the threshold).
- In 'window-strict' mode, the loop still breaks on the first failing
  window, returning the first window's adjustedComplexity.
- Per-ramp regressions in RampController and controller.score also stop
  returning 0 when minimum complexity is not sustained.
- scoreLowerBound (10th percentile of bootstrap resamples) can become
  non-zero on borderline runs where >= 10% of resamples previously
  returned 0, while median scores above the threshold are unchanged.

Change-Id: If236df53bb0a00f1008863cbe51798fcb8123f1f
Change-Id: Iea35400feddb704d54ef47b28d74adea4ad29bd8
High-level overview:
During scrolling in `#scrollSheet()`, the current Sheets subtest shifts
the existing `60vw x 60vh` canvas via `ctx.drawImage()` and sets a
clip for
the newly exposed 4px horizontal and vertical edges, but
`#drawBodyCells()`
and `#drawHeaders()` still iterate over all 238 cells in the sheet (190
body cells across 10 columns x 19 rows plus 48 header cells) on every
frame—allocating a `new Path2D()` clip and calling `ctx.measureText()`
and
`ctx.fillText()` for ~200 cells that lie completely outside the 4px
exposed scroll strip.

Why this is more realistic and less gameable:
Real canvas-based spreadsheet applications (such as Google Sheets)
viewport-cull cells against the visible or dirty scroll rect in
application code before measuring text or issuing draw commands, and use
direct `ctx.rect()` clips rather than allocating a `new Path2D()` object
per cell. Without application-level bounds culling in `#scrollSheet()`,
the subtest spends most of its frame budget executing JS cell iteration,
`Path2D` allocations, and `ctx.measureText()` / `ctx.fillText()` calls
for
cells completely outside the 4px dirty strip—measuring whether the 2D
canvas frontend performs early clip-bounds rejection rather than
measuring
visible spreadsheet scrolling and cell rasterization. Passing the
exposed
`dirtyRect` bounds into `#drawBodyCells()` and `#drawHeaders()` aligns
the
workload with real canvas spreadsheet engines while preserving the
original
`60vw x 60vh` sheet size and per-cell clipping behavior.

Summary of outcomes and changes:
- Redraw the exposed horizontal and vertical scroll strips in
  `#scrollSheet()` using separate rectangular `ctx.rect()` clips and
  pass
  each strip's `dirtyRect` bounds to `#drawBodyCells()`.
- Add spatial bounds culling to `#drawCells()`,
  `#drawColumnBackgrounds()`, `#drawGrid()`, `#drawHeaderRows()`, and
  `#drawHeaderColumns()` so cells outside the active viewport/dirty rect
  are not measured or drawn.
- Replace per-cell `new Path2D()` clip allocations with
  `ctx.beginPath(); ctx.rect(...); ctx.clip();` while preserving the
  existing `#drawCells()` FIXME comment and original `60vw x 60vh`
  canvas
  dimensions.

Benchmark impact (Chrome Android, Pixel 10 Pro XL, 60 Hz, N=2 paired
runs):
- Sheets score (at unchanged 60vw x 60vh canvas size): 7.88 [7.95, 7.81]
  -> 12.91 [13.13, 12.69] (+63.8%), driven by culling off-screen cell
  `measureText()`/`fillText()` and `Path2D` allocations outside the 4px
  scroll strip.

Change-Id: Ic9a854262697cd6792f0c7f284c9aa0de1ef0c64
@netlify

netlify Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for webkit-motionmark-preview ready!

Name Link
🔨 Latest commit d4707b5
🔍 Latest deploy log https://app.netlify.com/projects/webkit-motionmark-preview/deploys/6abe83b2eb00280008f7991b
😎 Deploy Preview https://deploy-preview-127--webkit-motionmark-preview.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.
🤖 Make changes Run an agent on this branch

To edit notification comments on pull requests, go to your Netlify project configuration.

@khokhlov962
khokhlov962 marked this pull request as draft October 1, 2026 16:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants