Files
mtgodot-poc/native_render/README.md
T
2026-09-27 19:44:41 -07:00

10 KiB
Raw Blame History

Native Vulkan renderer prototype

This is the standalone renderer path for the 40250 Render3DDraw and UIRenderCommand command streams, plus direct --live-client execution of the ported 40250 client (PythonBoot). SDL3 owns the window and input forwarding; Vulkan owns the swapchain (FIFO or IMMEDIATE/MAILBOX), VK_FORMAT_D32_SFLOAT depth buffer, VK_QUERY_TYPE_TIMESTAMP hardware GPU timer, key-indexed persistent vertex/index buffers, GPU skeletal skinning bone palette SSBO (set = 1, binding = 0), DDS/TGA/mem: glyph-page texture sampler descriptors, fixed-function 3D + 2D UI state pipelines, and draw submission. On macOS, the Vulkan loader uses MoltenVK over Metal. The existing Godot client remains the default playable path.

Build and run on macOS

Install Vulkan headers/loader, MoltenVK, SDL3 and glslc (shaderc). With Homebrew:

brew install vulkan-headers vulkan-loader molten-vk sdl3 shaderc

# 1. Standalone replay build (without port_platform)
cmake -S . -B build-native-render -DMTGODOT_BUILD_EXTENSION=OFF \
  -DMT_BUILD_NATIVE_RENDER_PROTOTYPE=ON -DCMAKE_PREFIX_PATH=/opt/homebrew
cmake --build build-native-render --target mt_native_render -j8

# 2. Full live-client build (links port_platform + mtpython + FakeLoginServer)
cmake -S . -B build -DMT_BUILD_NATIVE_RENDER_PROTOTYPE=ON -DCMAKE_PREFIX_PATH=/opt/homebrew
cmake --build build --target mtgodot mt_native_render port_fake_login_server -j8

# 3. Optimized Release (-O3) live-client build
cmake -S . -B build-release -DCMAKE_BUILD_TYPE=Release -DMT_BUILD_NATIVE_RENDER_PROTOTYPE=ON \
  -DMTGODOT_EMBED_PYTHON=ON -DCMAKE_PREFIX_PATH=/opt/homebrew
cmake --build build-release --target mt_native_render -j8

On macOS, mt_native_render automatically detects /opt/homebrew/etc/vulkan/icd.d/MoltenVK_icd.json (or /usr/local/etc/vulkan/icd.d/MoltenVK_icd.json) when VK_ICD_FILENAMES is not set in the environment. The build copies native.vert.spv, native.frag.spv, and, for the live client, python27.zip beside the executable. Keep these files together when moving the binary. In a macOS .app, put them in Contents/Resources (the directory returned by SDL's SDL_GetBasePath). MT_PYTHON_STDLIB can override the zip path.

Android integration

With an Android toolchain and SDL3 Android AAR/Prefab available to CMake, the same option builds libmain.so for SDLActivity. The mt_native_android_assets target prepares native.vert.spv, native.frag.spv, and python27.zip under <build>/native_render/android-assets (python27.zip is included when the embedded Python target is enabled); include that directory in the SDL app's assets.srcDirs. The app must allow network access for login. The 40250 Client directory must be placed at SDL_GetPrefPath("mtgodot", "native-render")/Client, with pack/Index present, before live mode starts. The Godot APK is a separate application and does not launch this SDL renderer.

The standalone arm64-v8a renderer cross-builds at Android API 24. The full live-client cross-build currently stops in extension/src/port/common/Win32Crt.cpp: the NDK exposes <iconv.h> at API 24 but does not declare iconv until API 28. This is a port-runtime prerequisite for an Android live-client APK; raising the minimum API level is not assumed here.

The Vulkan portability enumeration extension is selected only when advertised by the loader. DDS, TGA, and memory textures retain their existing decoders; JPEG, PNG, and BMP use the same portable decoder on macOS and Android. The vendored stb_image.h is upstream v2.30 (SHA-256 594c2fe35d49488b4382dbfaec8f98366defca819d916ac95becf3e75f4200b3).

Interactive Playable Modes

Run the full 40250 client interactively (infinite frame loop until window close, resizable SDL3 window with automatic Vulkan swapchain recreation and PythonBoot::SetUISize sync, full keyboard/IME text input, SDL hardware cursor built from the original cursor images, and SDL3 + AudioToolbox .wav/.mp3 audio):

# Interactive outdoor map session (auto-login via loopback FakeLoginServer)
./build-release/native_render/mt_native_render \
  --live-client "/path/to/40250/Server Client TMP4/Client" \
  --interactive --fake-mobs 24 --width 1280 --height 800

# Interactive login screen (stops at introLogin.LoginWindow for manual typing/login)
./build-release/native_render/mt_native_render \
  --live-client "/path/to/40250/Server Client TMP4/Client" \
  --login-screen --width 1024 --height 768

# Connect to an external 40250 Auth + Game server
./build-release/native_render/mt_native_render \
  --live-client "/path/to/40250/Server Client TMP4/Client" \
  --live-server 127.0.0.1:11002:13000 --login-screen

Synthetic benchmark

./build-release/native_render/mt_native_render --frames 60 --draws 64 --triangles-per-draw 333 --no-vsync

Timing metrics printed by mt_native_render:

  • p95_frame_ms, p99_frame_ms, max_frame_ms: wall-clock time for each measured update/render iteration, excluding startup and final GPU drain. These include vsync wait when enabled; use them alongside gpu_ms and the CPU breakdown.
  • game_update_ms: mean CPU time spent in PythonBoot::UIUpdate(), PythonBoot::UIRender(), and audio command draining per frame in --live-client mode.
  • prepare_ms: mean CPU draw-preparation time across all frames (including frame 0 cold-start geometry/texture uploads and pipeline creation).
  • steady_prepare_ms: mean CPU draw-preparation time on frames after frame 0 (bone palette copy, UI quad batching, command recording).
  • submit_ms: host CPU time around vkQueueSubmit (in FIFO mode this includes swapchain backpressure; pass --no-vsync to switch to IMMEDIATE/MAILBOX).
  • gpu_ms: true hardware GPU execution time between top-of-pipe vkCmdBeginRenderPass and bottom-of-pipe vkCmdEndRenderPass measured via VK_QUERY_TYPE_TIMESTAMP.

macOS real-server acceptance

Build the Release live-client target above, then start the evidence runner from the repository root:

# First verify the runner and renderer using the local fake server.
node script/native_mac_acceptance.mjs --fake --frames 180

# Use the real server's shared host, auth port, and game channel port.
node script/native_mac_acceptance.mjs --server HOST:AUTH_PORT:GAME_PORT

Real-server mode checks both ports before starting, opens the native login screen, and collects a redacted client log, one-second process RSS samples, and a JSON report under build/native-acceptance/. Enter credentials in the app, then exercise login, character selection, movement, combat, map changes, inventory, chat, window resize/focus, and visual comparison with the current client. Close the window after at least 30 minutes. The report records whether that minimum was met, but remains NEEDS_MANUAL_REVIEW until those actions and visual results are checked by a person. It does not assert that RSS alone proves no GPU leak.

--live-server currently accepts one shared host for the auth and game ports. If those endpoints use different hosts, update the native connection setup before claiming a real-server pass. Credentials are entered in the client UI; the runner never puts them in arguments or the report.

Run the real 40250 client benchmark in native Vulkan (--live-client)

When built with port_platform, mt_native_render boots system.py, logs in via loopback FakeLoginServer, enters the outdoor map with 1..64 monsters, renders the complete 3D scene (40250 hardware-transform terrain splats, animated water patches, gradient skybox & scrolling clouds, SpeedTree forest bark/leaf geometry, GPU-skinned characters with stage-1 specular sphere-maps, and 2D UI/minimap/text-tails/software cursor), and optionally writes a Version 5 .mtdr capture:

./build-release/native_render/mt_native_render \
  --live-client "/path/to/40250/Server Client TMP4/Client" \
  --fake-mobs 64 --frames 180 --gpu-skinning --no-vsync \
  --capture-out /tmp/mt_full_64mobs.mtdr

Compare against CPU skinning (GrannyDeformVertices on CPU + per-frame vertex buffer re-uploads) with --no-gpu-skinning:

./build-release/native_render/mt_native_render \
  --live-client "/path/to/40250/Server Client TMP4/Client" \
  --fake-mobs 64 --frames 180 --no-gpu-skinning --no-vsync

Replay a captured frame (.mtdr v1 / v2 / v3 / v4 / v5)

./build-release/native_render/mt_native_render \
  --capture /tmp/mt_full_64mobs.mtdr --frames 180 --animate-bones --no-vsync
  • Pass --animate-bones to animate the bone palette SSBO each frame without re-uploading any vertex buffers (uploads stays equal to unique static geometries uploaded on frame 0).
  • Pass --animate-first-draw to increment the first draw's geometry_revision each frame after frame 0 and verify incremental GPU buffer re-uploads.

Capture format Version 5 (backward-compatible with Versions 1–4) stores:

  1. 3D draws (Render3DDraw): geometry_key, geometry_revision, matrices, positions, normals, UVs, diffuse colors, indices, D3D8 fixed-function states, texture0 / texture1 names, and GPU skinning data (bone_indices, bone_weights, bone_matrices).
  2. Self-contained textures: .dds, .tga, .jpg, .png, and .bmp pack bytes plus "MTRA" raw RGBA memory textures (mem:<id>@<revision> font glyph pages).
  3. 2D UI stream (UIRenderCommand): canvas size (ui_width, ui_height) and all Bar, GradientBar, Line, and Image commands (including behind_3d, clip rects, minimap mask UV coordinates, and software mouse cursor quads).
  4. Per-draw fog color, vertex/table mode, range flag, start/end distance and density. Older captures omit these fields and replay without fog.

The native shader now evaluates recorded D3D8 stage 0/1 color and alpha operations, including the original cloud operation (D3DTOP_MODULATEINVALPHA_ADDCOLOR = 20), and applies the captured linear or exponential fog. Expanded UI image modes use the 40250 blend factors for screen/color-dodge and modulate. These state fixes do not by themselves establish pixel parity with a Windows 40250 screenshot; compare the same map, time, camera and UI state before treating a color difference as resolved.