Disclaimer: As always, feel free to correct me, as it might me help understand better and maybe fix some bugs. I'm still exploring the hardware, I might miss something or say something plain wrong. The following is not source of truth value, only my current understanding of things applied to the port.
Tethys Odysseus - Oddworld: Abe's Oddysee on Sega Saturn - update (11-17 August)
── LIGHTS: THE VDP2 DOES THE BLENDING ────────────────────
A VDP1 sprite cannot blend against a VDP2 background. VDP1 half-transparency (colour calculation mode 3) works on "the pixel data of the original graphic and the pixel data read from the write coordinates" - the frame buffer, which holds VDP1's own output only - and the manual notes it takes six times longer than drawing with no colour calculation. It also wants RGB codes; our cels are paletted.
So VDP2 adds instead. CCCTL (1800ECH) bit 8, CCMD = 1, is "add as is (additive blending)": saturating addition of the sprite onto the CAM, which is PSX blend mode 1 (B + F) reproduced 1:1. slColorCalc(CC_ADD | CC_TOP). Register only, no bank, no fill.
The per-sprite gate is NOT the CC bits, which only pick a ratio slot. VDP2 tests the sprite's PRIORITY field against a colour calculation condition number under one of four conditions (SPCTL 1800E0H bits 13-12: PR<=CCN, PR=CCN, PR>=CCN, or colour-data-MSB). PR=CCN with the number 1; every ordinary sprite ships PR = 0, so only tagged sprites blend. Sprite type 3 (SGL's boot SPCTL 0x0023) is MSB-first SD:1 | PR:2 | CC:2 | DC:11, so PR=1 is bit 13, CC=1 is bit 11, and tagging is COLR |= 0x2800.
8bpp only, and that matters: CMDCOLR is colour bits only in a colour BANK mode. In CL16Look it is a VRAM address for the lookup table, in CL32KRGB it is literal RGB555 - OR-ing 0x2800 into either corrupts a palette pointer or the red channel. Hence five files promoted 4bpp -> 8bpp offline in the converter (SQBSMK, EXPLO2, SPLINE, FLNTGLOW, BLOOD), 11.6 KiB of a 434 KiB texture heap.
The "this is a light" discriminator is AO's own blend mode: eBlend_3 is authored by exactly six classes (DoorLight.cpp:111, DoorFlame:42, MotionDetector:363, ZapLine:46, Particle:55, MainMenu:4012), all light or energy effects. eBlend_1 takes the tag on a second predicate rather than by widening the first, which would have put a halo on every explosion shard.
One rule that cost a build: mesh is a SUBSTITUTE for blending, not a companion to it. Stippling prims that were already blending is what I saw as a black checkerboard in Abe's green gas - VDP1 skips meshed pixels entirely, so the CAM shows through. The gate is now the cel's ability to blend (depth != 8).
Three claims from last update are withdrawn: the fart cloud is SQBSMK.BAN and is gated on the fart-gas cheat, not ABEGAS.BAN; BLOOD.BAN does not blend on PSX either (its STP indices are 3..15, all pure white, used zero times); and "eBlend_0 + semiTrans" describes ~330 AnimIds in R1, so it identifies nothing.
── THE SPRITE CEILING: 60 -> 128 ─────────────────────────
The cap came from SGL_MAX_POLYGONS = 64, carried over from a sibling project's note ("MaxPolygons > ~64-80 starves the TLSF heap -> boot loop"). That was measured against a pool of ~80 KB. It is a trade, not a hardware limit - the work area comes 1:1 out of the TLSF pool - and the path reclaim below took the pool to 158 KB. So: SGL_MAX_POLYGONS 64 -> 192 (work area 0x36F8 of its 0x4000 reservation, 2,312 B spare), cap 60 -> 128, kZBase 80 -> 148.
The cap is measured every frame, not asserted, because crossing SGL's sort-table bound is not a lost sprite - it is command bytes written inside the sort table, i.e. a silent SH-2 reset.
SGL fills three of those terms in at slInitSystem, so no map read settles them. If 192 polygons pay for 128 sprites the answer is 128; if not, the build degrades instead of resetting. Rank-based shedding is now inert.
── THE ELEVATOR CHAIN ────────────────────────────────────
Drawing wrong since S5, four builds of gauges spent on the assumption that something was dropping it. The gauges finally said 30 rope rects offered, 30 drawn, nothing lost - so I read the artefacts instead. The shipped R1ROPES.BAN texture, re-decoded off the disc, is a solid 6x16 opaque block; the C07 background has no chain painted in it. The chain was on screen and wrong, not absent.
Cause: the far-plane decimation gate compared destination height against the slot's source height and never asked how many SOURCE rows the draw sampled. Its own comment names the discriminator it failed to use - AO shrinks the QUAD, not the UVs - but a vertically CLIPPED poly shrinks the quad and its UVs together. Rope::VRender clips its end segments, one clipped by 3+ rows trips the gate, and decimation rewrites the SHARED slot in place (all 14 segments point at one Animation, Rope.cpp:137). One clipped end segment halved the whole chain.
Then it was too large, then too thin, and both were my mistakes too: removing a max(8) content floor stopped column DUPLICATION but exposed point sampling, which cannot represent a feature narrower than its own step. The rope is 3 opaque columns of a 4-wide cel. The prescaler now asks COVERAGE - if any covered source column is opaque, the destination is opaque and takes the blend of the covered samples. Verified by re-reading the shipped pack: 2 opaque columns, against 6 and 1 in the two builds before. Cost is printed permanently on every converter run: 122,113 texels rescued, 7.69% of the opaque census.
── 99 PATHS FOR A GAME WHOSE BIGGEST LEVEL HAS 21 ────────
AO's 46 per-level path arrays are declared [kMaxPaths] and land in HWRAM .data rather than .bss because their initialisers are non-zero: 97.1 KB in the map file, on a machine whose entire ao_new pool was 82 KB. kMaxPaths is 99 only because RELIVE stubbed Path_Get_Num_Paths to return it as a blanket bound - the real line is still there, commented out. Largest tables in the whole game are R1's and R2's at 21 entries; gMapData's field_18_num_paths agrees (20 R2, 11 D2, 9 F2, 6 R6).
99 -> 24. Pool 82.1 -> 159 KB, image 91,728 B smaller, and four flip-path loops stop walking 99 entries to look at 21. Safe because PathData and CollisionInfo are never indexed by a runtime path number - the PathBlyRec initialisers take their addresses at compile time, all 287 checked, highest is 20 - and the one runtime index now carries a named fatal.
Transferable bit: one non-zero initialiser moves a whole array from .bss to .data, i.e. out of free LWRAM and into the image. Run objdump -h before hunting a big culprit.
── LZ4 BACKGROUNDS, AND WHAT THE DRIVE ACTUALLY COSTS ────
.CAM backgrounds ship through an LZ4 container (BE raw length, then length-prefixed independent 8 KiB blocks). The converter decompresses every file it emits and refuses any that does not come back byte-identical. 49 R1 cameras, both languages: 4,779,860 -> 3,279,386 B (1.458x); R1.LVL 7,264,256 -> 5,761,024; largest record 165,364 -> 108,181. Cart slots 176,128 -> 131,072 B, 8 -> 12 resident screens, free.
LZ4 over deflate (1.701x) because the figure of merit is ratio MINUS decode, and deflate wants bit-wise Huffman on an SH-2 with no barrel shifter. The decoder has a longword fast path for literals and a deliberately byte-wise path for matches, since LZ4 matches overlap their own output; it typedefs its pointer width so the same source compiles for a 64-bit host and was verified against the shipped R1.LVL.
With the cache serving nearly every flip, the screen change became legible: the read is flat at 179-229 ms while the OPEN is 507-1364 ms, 73-86% of the whole change. Cost does not follow bytes - 16 KB costs 415 ms, 66 KB costs 1118 - which fits ~190 ms fixed plus ~71 KB/s, half what the same drive gives on large reads. Per-read overhead, not throughput.
── THE IN-PLAY SLOWDOWN, MEASURED ────────────────────────
From the capture in the video: full debug overlay, 3 min 43, sampled every 2 s, 111 frames, every row transcribed.
- Not accumulation. HWRAM free is HIGHER at 3:20 than at 1:00, LWRAM constant, and no capacity gauge is ever non-zero across all 111 samples. Nothing is dropped - frames are late.
- Fully reversible: returning to the opening screen restores exactly 30.0 fps.
- It follows CONTENT: 30.0 fps at 13-15 sprites, 14.3-16.7 fps at 42-64. Marginal cost ~1.2 ms, roughly 32,000 SH-2 cycles, per additional sprite.
Clock calibration, since cycle figures are worthless without it: the tick source is the free-running timer at PHI/128, so 1 raw tick = 128 SH-2 cycles and 208 ticks = 1 ms - 26,624 cycles/ms in 320-wide NTSC. NOT 28,636, which is the 352 clock; anything priced against that is 7.4% high.
One law I cannot explain yet: the ordering-table walk costs ~0.38 ms per SUBMITTED SPRITE and does not track prim count. 419 prims / 15 ms on one screen, 309 prims / 25 ms on another. Fewer prims, more time.
Which phase dominates a typical frame is deliberately not answered here: the overlay row I would have to quote is a per-screen MAXIMUM, and I do not yet trust myself to read a maximum as a mean. Next update.
── THE LCD MARQUEE IS 20% OF A FRAME ─────────────────────
With the marquee running, prim count, upload count, byte count and textured-rect count are all flat while the rate goes 24 -> 30 fps the instant the text finishes. Nothing else on the row moves.
The named cost is structural: the glyph atlas is read from VDP1 VRAM at 0x25C00000, the cache-through partition, so every source halfword is a real B-bus round trip that cannot be cached and contends with VDP1's own drawing. A scrolling panel re-composites every frame by construction. Lever: shadow the atlas in a cacheable bank.
Footnote: the blit has been bracketed and timed for dozens of builds, paying two clock reads a frame, and printed on no row.
@slygamer had to ask why the LCD is expensive before I looked at my own instrument.
── EXODDUS, AND WHY IT IS ON A DIFFERENT BRANCH ──────────
Abe's Exoddus now reaches gameplay on the Saturn, on the same renderer, sound path and CD path as Oddysee.
It is built from RELIVE's beta branch rather than the master our Oddysee port sits on, for a simple reason: on master the Exoddus game code still carries 70 NOT_IMPLEMENTED sites, and on beta it carries none. Beta is 1,783 commits ahead. That is not a free upgrade - it brings JSON parsing at runtime, std::mutex and a thread pool, and it reworks the renderer seam from 28 virtuals down to 12 - so Oddysee is NOT migrating. Exoddus lives on its own branch in its own worktree, and the two games share the Saturn platform layer instead of sharing an engine revision.
The memory wall closed without needing the RAM cart. Biggest single lever was three lines of desktop code: <iostream> arriving through a logger header and an .ini writer pulled 586 KB of C++ stream machinery. Second was defining all fifteen of libstdc++'s std::__throw_* symbols plus the terminate handler - a static archive is pulled in one MEMBER at a time, so defining some of them still drags the member in for the rest, and defining all of them means the exception machinery, unwinder and name demangler included, is never linked at all. -95,808 bytes of text, in a program compiled -fno-exceptions.
What actually stopped it booting was not size, though. SGL's preloader runs before main and creates the malloc heap as (work area - end of .bss). Once .bss overran the window that subtraction went NEGATIVE and was passed as an unsigned size, so TLSF wrote its control block past the top of HWRAM - which mirrors back down onto our own .text. Dead before main, and it would have died at 900 KB exactly as surely as at 3 MB. Fixed by a linker map that puts .bss and the sbrk arena in LWRAM: .bss is NOLOAD, so it is the only section that can move without growing the binary.
── TWO THINGS I GOT WRONG ────────────────────────────────
1. The predictive prefetch. Speculate on the neighbouring screen's background while the player walks, serve the flip from RAM. On Ymir it did everything promised: the heavy flip 440 -> 200 ms, bytes read at the flip 70 -> 10 KB. On hardware
@slygamer reported it slower on BOTH axes - longer loads and lower frame rate - against the same build with it off. Removed entirely. The precedent was already on the wall, which is why it shipped default-off behind a chord: a sibling project (Mimas, for a change) measured this exact shape, a persistent file handle kept busy, as a hardware regression on an SD ODE. Ymir could only ever confirm the mechanism RAN, never that it PAID. The captures did pay for the drive finding above.
2. The particle spawn cap. Halving blood and smoke per impact looked obvious. Putting the toggle state on the overlay row - the missing witness - made the A/B valid for the first time: frame rate 20/20/22/23 with the cap off, 22/20/23/24 with it on. Flat, if anything worse, while upload volume genuinely halves. Removed with its chord rather than left inert. The particles were never a speed problem, they were a sprite-count problem: one bullet asking for 20 of 60 slots, which is what the ceiling work fixed.
── STILL BROKEN, NAMED ───────────────────────────────────
- Frame rate: 23.7 fps average in-play, 14.3 worst, on screens carrying 40+ sprites.
- Loading: 22 screen changes in the capture, 40.82 s of black, mean 1.86 s, 17.7% of a 230.7 s run. Per-read overhead, per above.
- A white block replaces the barrel animation on some screens, visible in the video, and it is ours. Two doors of this bug are shut - a lamp pinning a CRAM bank for a whole run, and the per-screen VDP1 heap rewind handing the lamp's texture addresses to the next screen's uploads because those textures sit outside the slot table the rewind walks. Both fixes are in this build. Still there. Third instance, not found.
- Lamp halos are too wide. The originals are 8-12 pixels across; I widened ours for legibility and a fidelity pass is owed.
- New, and unexplained: a floor hatch that is orange on PSX renders VIOLET here. Measured on the raw captures, same room, same object - PSX hue 17 deg, RGB (112, 35, 26); Saturn hue 278 deg, RGB (49, 0, 76) at 0.97 saturation, green channel essentially zero. Looks much more like a palette entry taken from the wrong place than a conversion rounding issue, same family as two CRAM classification bugs I have already fixed, but unproven.
-
@Wesker's hanging note, not audible in the video, that happens on mudokon save and secret room reveal: AO factored its note-on into a function whose signature carries neither the sequence index nor the MIDI channel, so it never writes the tag a sequence stop scans to decide which voices it owns. AE writes it, AO never has. Inaudible on PSX because the hardware kills the voice with a second key-off; here, 49 of the sound bank's live tones carry a release rate that works out to a 32.7 s quadratic ramp, so a released note is still at 99.2% volume three seconds later. Fixed the day after this capture.
── SINCE THIS CAPTURE ────────────────────────────────────
Three days of work did not make the video. Briefly, for the next update:
- All fourteen levels are on the disc now - 249.2 MB of PC data down to 77.3 MB of packs
- Nine of the 115 object factories demand a resource their own load branch never fetches, and five have no load branch at all. That is one design, not nine bugs, and hardening the check into a fatal was the wrong call.
-
@Wesker's hanging note is fixed.
- Exoddus keeps moving: solid foreground, text, palette overrides, and a screen change split into drive time versus everything else.
── WITH THANKS ───────────────────────────────────────────
@slygamer: dozens of builds, dozens of captures, and the patience to run the same route every time so two builds are actually comparable. The capture carrying the whole frame-rate section is hers, and she killed the prefetch.
@Wesker: found the hanging note, which turned out to be a real PSX-vs-Saturn asymmetry in how a sequence stop reaches a voice.
@carlos24_: for this original post and port idea