KINDLEFALL Play free
Part 9

Hundreds of things on screen, and no stutter

How the game stays smooth on our test machine, with before-and-after Chrome profiles of old builds.

Kindlefall fights get busy. A dozen elites with auras and ward bubbles, a Pyromancer throwing arrows and fireballs, molten creatures spitting lobs, particles everywhere, lava moving under all of it. On our test machine it stays smooth through all of that. This post is how, with the measurements to back it up, including before-and-after profiles of old builds checked out from git.

The stress sceneTwelve elites (molten crawlers and Cinderlings, smiths, a golem, warriors, cultists and mages, each with an elite trait) around the Pyromancer on Floor IV, with boons stacked. This is the horde used in the profiles below.

The budget

The game caps itself at 60 frames a second by default, which gives every frame 16.7 ms. The cap exists because the development MacBook ran hot early on: its display runs at 120 Hz, and running uncapped meant doing twice the GPU work for no visible gain. Menus and the title run at 30 fps, and the High preset renders at 1.5 times the CSS resolution instead of full Retina 2x (Ultra is there if you want it).

The question is never just "is the average frame fast". It's "does any frame ever take longer than 16.7 ms", because one slow frame is a visible hitch. So every measurement here reports percentiles and the worst frame, not just averages.

The same fight with the in-game F3 performance overlay in the corner: 59 fps, 17.09 ms between frames, 110 draw calls, render scale 1.00, quality High, 731 particles, 9 enemies
The F3 overlayPress F3 in game to see it. Note that "17.09 ms" here is the time between frames at the 60 fps cap, not the time a frame costs; the cost is in the charts below. Captured in real time during the horde scene.

What a frame actually costs

In Sprint #2 a dedicated Performance team profiled the game with real GPU rendering (Metal), forcing the GPU to finish every frame so its time was counted, with CPU and allocation profiles, and a tracker for shaders compiled late. Their scenarios: the busiest fight on every floor, every boss with its phases forced, a fight with four rounds of boons stacked, and a twelve-elite horde.

The finding: on Apple Silicon a frame costs about 1.7 ms of CPU and GPU, around a tenth of the budget. Fights on all seven floors measured a median of 1.6 to 1.8 ms and a 95th percentile of 2.0 to 2.4 ms after their work; the boon-stacked fight 2.1 ms median and the horde 2.3 ms. The game was never short of raw speed. What players feel are hitches: single long frames. The team found five causes and fixed all of them:

  1. Crumbling floors on Floor VI. Every tile that dropped rebuilt the enemies' navigation grid against every wall and pit on the whole floor: a 20 to 30 ms frame per tile, mid-fight. The rebuild now works in small tiles and only looks at obstacles nearby. A tile drop went from about 25 ms to under 1 ms, and building every floor's grid at load from 1.4 seconds to about 70 ms.
  2. Shaders compiled on first use. A handful of effects (Floor IV's trip-hammer beams, Floor V's surges and breakers, the Storm boons' bolts, the aim line and a few more) compiled their shader the first time they appeared, mid-fight: 13 to 450 ms once, and one 2-second outlier. They are now compiled during the warm-up behind the loading screen.
  3. Adaptive resolution churn. The game lowers its render resolution when frames get slow. Each change reallocates the render targets, 20 to 110 ms, and the old rule changed it too eagerly. Now it steps down only after 1.5 seconds of slow frames, steps up only after 6 seconds of headroom (and 20 seconds after a drop), and only in a calm room unless frames are badly slow. If frames stay slow even at reduced resolution, the game steps down one quality preset for the session, with a toast, and never saves it.
  4. Shadow-pass churn. One shared shadow material made three.js re-derive its shader program about 13 times a frame, creating garbage every frame. Skinned and instanced shadow casters now get their own material.
  5. Loading. Decoding the sound-effect sprite (1 to 2 seconds) ran one after the other with building the floors. Now they overlap, and the warm-up renders one floor per frame behind the loading screen instead of one long block. Load to title went from about 4 seconds to about 2.

Bosses on all seven floors went from worst frames of 4 to 24 ms (with shader-compile spikes of 13 ms to 2 seconds in some runs) to 3 to 9 ms, with no compiles mid-fight.

Measured onThe Performance team's figures above: MacBook Pro, Apple M3 Max, Chrome with Metal, High quality, 1280×720 at 2x pixel density, uncapped frame rate, the GPU forced to finish every frame. The machine was also running other teams' tests at the time (load average 30 to 70), so medians are noisy; worst frames and hitches are what to compare.

Before and after, profiled

To show the difference rather than just describe it, we checked three builds out of git into separate folders, served each on its own port, and ran the exact same scripted scenes on each in GPU Chrome, recording a real Chrome performance trace:

The scenes: the twelve-elite horde above, the Crumbling Gallery on Floor VI (where tiles drop), the Storm Seraph fight, and arriving on Floor V. Each ran twice per build, alternating builds. The test bot played; the hero was invincible so every run lasted the same time. "Frame cost" is the time from the start of a frame to the GPU finishing it.

Crumbling Gallery (Floor VI): frame cost over 10 seconds 0 2 4 6 8 10 12 14 0s 2s 4s 6s 8s 10s ms per frame (CPU + GPU) Before Sprint #2 (b1ccb70) After perf round 1 (2d02b95) Today (12eb2f9)
The Crumbling GalleryFloor VI, tiles dropping under the fight. The old build's spikes are the navigation rebuild on each dropped tile; the fixed builds stay flat. Second run of each build shown.
Horde of 12 elites (Floor IV): frame cost over 12 seconds 0 2 4 6 8 10 0s 2s 4s 6s 8s 10s 12s ms per frame (CPU + GPU) Before Sprint #2 (b1ccb70) After perf round 1 (2d02b95) Today (12eb2f9)
The hordeOn this machine all three builds were already far inside the 16.7 ms budget here: the horde was never the problem. Second run of each build shown.
SceneBuildp50 msp95 msp99 msworst mslong tasksdraw calls
Horde: 12 elites, Floor IVBefore2.83.94.67.90142
Round 12.73.94.58.80139
Today2.94.14.97.90143
Crumbling Gallery, Floor VIBefore2.03.57.612.2055
Round 12.03.13.65.3058
Today2.03.43.86.0052
Storm Seraph fightBefore1.72.52.94.5060
Round 11.82.83.43.9059
Today2.03.33.84.7060
Arriving on Floor VBefore2.03.43.78.7063
Round 12.23.54.410.4065
Today1.93.13.78.8065

Before = commit b1ccb70, Round 1 = commit 2d02b95, Today = commit 12eb2f9. Frame cost in milliseconds, both runs pooled; draw calls are the most seen in one frame.

Load to title (seconds, mean of two runs) Before Sprint #2 (b1ccb70) 4.9 s After perf round 1 (2d02b95) 2.9 s Today (12eb2f9) 3.2 s
Load to titleFrom navigation to the title screen, at the same High quality. The machine was busier during these runs than during the Performance team's, which measured about 2 seconds after the fix.
JS heap during the horde fight (MB) 0 60 120 180 240 300 0s 2s 4s 6s 8s 10s 12s MB in use Before Sprint #2 (b1ccb70) After perf round 1 (2d02b95) Today (12eb2f9)
Memory during the hordeThe JavaScript heap rises and falls as the garbage collector runs; what matters is that it comes back down to the same floor. Garbage-collection time inside the traced 12 seconds: Before Sprint #2 (b1ccb70): 31.9 and 9.0 ms; After perf round 1 (2d02b95): 32.4 and 17.6 ms; Today (12eb2f9): 16.3 and 16.6 ms.

The honest summary: on this machine every build was already well inside the budget in these scenes, and none of the runs recorded a single long task (a main-thread block over 50 ms). The Sprint #2 work shows up where the team said it would: in load time, and in the worst frames of the room where the floor falls away. The rest of their fixes target things these scenes don't hit on a fast Mac: first-use shader compiles in specific effects, resolution changes on slower machines, and memory over a long session.

The traces themselves are here, so anyone can open them in Chrome DevTools (Performance panel, "Load profile"): before, after perf round 1, today (horde scene, about 4 MB each).

Measured onThese profiles: MacBook Pro, Apple M3 Max (48 GB), Chrome 154.0.8037.93 with Metal (driven by Playwright), High quality, 60 fps cap, a 1280×720 window at 2x pixel density. Other work was running on the machine at the same time (load average 40 to 63), so treat small differences as noise.

Lots on screen, few draw calls

The other half of performance is not doing work in the first place:

Memory, and power in menus

A long session is a different test from a single fight. The Performance team ran ten restarts, ten Save & Quit and Continue cycles, ten rounds of every overlay, four full descents and ten 30-second Gauntlet sessions, forcing garbage collection before every measurement, and found four leaks: skeleton textures that were never freed, procedural monsters' geometry, summoned enemies orphaned by a list being modified while it was walked, and Photo Mode creating a new shader every time it opened. After the fixes, restarts, descents and overlays are flat; textures on the title went from 59 to 32. Part 7 has the full table.

Memory for sound got the same care: music loads per floor and frees the old floor's (all of it decoded at once was about 400 MB), and the sound-effect sprite is decoded in mono at 32 kHz, about 110 MB instead of 325. Those two figures are calculated from the audio lengths and formats, not measured on a particular machine.

Paused screens don't need 30 redraws a second of a frozen world. The pause menu, Settings, the Codex, Trophies and Vows now redraw it every third frame while still reading input every frame: busy time per second fell from 32 ms to 13 in the pause menu, 30 to 12 in Settings and 22 to 8 in the Codex (MacBook Pro M3 Max, Chrome, High quality, 1512×945 at 2x).

The referee

None of this would hold without the test battery. Before every ship, the bot clears all seven floors, and a regression run records frame-time percentiles and draw calls for each one. At the overnight sprint's hand-back, that run showed every floor cleared, a 95th-percentile frame interval of 17.3 to 18.2 ms against the 60 fps cap (in other words, frames arriving on time), at most 64 to 80 draw calls, no errors, and a worst frame of 9.6 ms on any floor arrival, down from 74 ms (headless Chrome with Metal on the same MacBook). A change that makes the game slower shows up in a number before it shows up in someone's hands.

Play Kindlefall

In your browser, with a keyboard and mouse or a controller.