完善客户端功能并加入40250一致性审计

This commit is contained in:
shen
2026-09-21 16:38:58 -07:00
parent 3eb001a1fd
commit 1522a8a1a1
22 changed files with 876 additions and 105 deletions
@@ -0,0 +1,77 @@
---
name: metin2-40250-parity-audit
description: Audit, implement, fix, or verify behavioral parity between the Metin2 40250 Windows C++ client and this Godot client. Use whenever a request mentions 40250 comparison, 1:1 parity, missing original-client behavior, parity regressions, or continuing the client audit. Do not use for unrelated features that have no 40250 compatibility requirement.
---
# Metin2 40250 parity audit
Treat the reachable 40250 implementation semantics and its observable behavior as the reference contract. The Godot architecture and APIs may differ, but its gameplay algorithms, branch conditions, state-transition order, constants, units, timing/event sources, resource/data sources, protocol side effects, and failure/cleanup behavior must not differ unless the difference is a documented platform adapter with evidence of semantic equivalence. Compare complete call chains, not filenames or similarly named functions.
## Persistent state
The audit ledger is the source of truth:
- `audit/manifest.json`: current contract states and evidence links.
- `audit/history.jsonl`: append-only state-change history.
- `audit/contracts/`: detailed evidence for individual contracts.
- `audit/reports/`: generated summaries; never treat these as the source of truth.
Before auditing, fixing, or claiming parity:
1. Run `python3 .agents/skills/metin2-40250-parity-audit/scripts/audit_ledger.py refresh --write`.
2. Run `python3 .agents/skills/metin2-40250-parity-audit/scripts/audit_ledger.py report`.
3. Read the relevant existing contract and tests. Do not repeat a still-valid `TEST_VERIFIED` audit unless the user explicitly requests revalidation.
4. If no contract exists, add one with a stable behavior ID before marking work complete.
Read [references/audit-schema.md](references/audit-schema.md) whenever creating or changing ledger entries. Read [references/project-map.md](references/project-map.md) when locating reference code, current implementation, existing gap documents, or test runners.
## Audit method
For each behavior:
1. Define the external trigger, preconditions, state transitions, outputs, timing, interruption, failure, and cleanup behavior.
2. Trace the complete reachable 40250 call chain, including resource-driven branches and compile-time feature flags.
3. Build a branch-by-branch equivalence table mapping the reference preconditions, decisions, formulas, state writes, ordering, timing sources, resource reads, outputs, and cleanup paths to the current native extension, GDScript, resources, protocol handling, and UI.
4. Classify every difference as missing, partial, wrong order, wrong value/unit, wrong algorithm, wrong timing/event source, wrong resource/data source, duplicate, platform adapter, or intentionally excluded.
5. Treat automated output equality as necessary but not sufficient: tests can miss branches, so a different algorithm cannot be approved merely because sampled outputs currently match.
6. Fix root mechanisms rather than coordinates, entity IDs, individual assets, one-off timing constants, or simplified approximations.
7. Add or strengthen an automated test that fails before the fix and covers the reference behavior. Include boundary, rejection, interruption, and cleanup paths when they materially affect the behavior.
8. Run the narrow test first, then related subsystem tests, then `git diff --check`.
9. Update the contract's implementation-equivalence matrix, manifest, fingerprints, and history in the same change. Never mark `STATIC_VERIFIED` or `TEST_VERIFIED` without complete equivalence evidence.
## Implementation-equivalence gate
Source text and engine-facing APIs do not need to be identical, but the implementation must be semantically unified with 40250. Before verification, prove all of these independently:
- identical effective preconditions and early-return rules;
- identical reachable branch structure and branch outcomes;
- equivalent algorithms and formulas, without simplified substitutes;
- identical state mutations and mutation order;
- identical constants, tolerances, coordinate conversions, and units after explicit platform conversion;
- identical timing authority and event source, such as `.msa` events rather than replacement timers;
- identical resource, table, and protocol data authority rather than hardcoded substitutes;
- identical network and externally visible side effects and their ordering;
- identical interruption, rejection, rollback, failure, and cleanup behavior.
Permitted differences are limited to documented platform adapters such as C++ containers to Godot collections, DirectX matrices to `Transform3D`, or Windows input APIs to Godot input APIs. For every adapter, document both sides, the conversion invariant, and a focused equivalence test. If any material item differs or lacks proof, status must remain `PARTIAL` or `MAPPED`.
## Evidence rules
- `MAPPED` means only that both sides were located.
- `STATIC_VERIFIED` requires a documented call-chain and implementation-equivalence matrix covering every material branch. Similar output alone is insufficient.
- `TEST_VERIFIED` additionally requires meaningful automated behavior tests with a recorded passing result; tests do not waive the implementation-equivalence gate.
- Engine/platform replacement may be `EXCLUDED` only when the replacement and externally observable verification are documented.
- A source or test fingerprint change makes prior verification `STALE`; investigate only the affected contracts.
- Existing prose in `docs/CLIENT-GAP.md` and similar documents is useful evidence, but it is not a current verification state unless represented in the ledger.
- Do not claim that code inspection proves visual, timing, input-feel, driver, or Windows-specific parity. Record those limits explicitly.
## Scope and safety
- Preserve unrelated dirty-worktree changes.
- Prefer the active files from the 40250 Visual Studio build; do not audit disabled, third-party, or obsolete code as product behavior without evidence that it is reachable.
- Keep credentials, live-server data, copyrighted binary dependencies, generated captures, and local run outputs out of the ledger.
- Do not change live servers or external systems unless the user requested it.
## Completion report
Report the contract IDs changed, reference and implementation call chains, discrepancies fixed, tests run, remaining unverified branches, and resulting ledger status. A subsystem is complete only when its in-scope contracts have no unexplained `UNMAPPED`, `PARTIAL`, `STALE`, or `REGRESSION` entries.
@@ -0,0 +1,108 @@
# Audit ledger schema
`audit/manifest.json` is machine-readable and version controlled. Keep one entry per externally meaningful behavior, not one entry per source file.
## Contract fields
Required fields:
- `id`: stable dotted ID such as `combat.local.normal_attack`.
- `title`: concise behavior name.
- `subsystem`: `network`, `lifecycle`, `movement`, `combat`, `skill`, `world`, `item`, `ui`, `resource`, `render`, `audio`, or another stable domain.
- `priority`: `P0`, `P1`, `P2`, or `P3`.
- `status`: one of the statuses below.
- `reference.files`: paths relative to `manifest.reference_root`.
- `reference.symbols`: relevant 40250 symbols.
- `implementation.files`: paths relative to the repository root.
- `implementation.symbols`: corresponding native or GDScript symbols.
- `equivalence`: implementation-equivalence matrix described below.
- `platform_adaptations`: documented engine/platform substitutions; use an empty array when none exist.
- `evidence.contract`: detailed Markdown contract path relative to the repository root.
- `evidence.tests`: test paths relative to the repository root.
- `evidence.last_test_result`: `PASS`, `FAIL`, or `NOT_RUN`.
- `remaining`: explicit unverified branches; use an empty array only when none remain.
Optional generated fields:
- `fingerprints.reference`: combined SHA-256 for the listed reference files.
- `fingerprints.implementation`: combined SHA-256 for implementation files.
- `fingerprints.tests`: combined SHA-256 for test files.
- `stale_reasons`: generated reasons for invalidation.
- `verified_commit`: Git commit at the last completed verification.
- `notes`: concise information that does not belong in the detailed contract.
## Implementation-equivalence matrix
Every contract that reaches `STATIC_VERIFIED` or `TEST_VERIFIED` must contain all fields below with the value `VERIFIED`:
```json
{
"equivalence": {
"preconditions": "VERIFIED",
"branch_structure": "VERIFIED",
"algorithms_formulas": "VERIFIED",
"state_transition_order": "VERIFIED",
"constants_units": "VERIFIED",
"timing_event_sources": "VERIFIED",
"resource_data_sources": "VERIFIED",
"protocol_side_effects": "VERIFIED",
"interruption_failure_cleanup": "VERIFIED"
}
}
```
`VERIFIED` means the detailed contract contains a branch-by-branch comparison and no material semantic difference. Matching a few outputs or passing only happy-path tests is not enough.
Language and engine boundary substitutions belong in `platform_adaptations`:
```json
{
"platform_adaptations": [
{
"reference": "D3DXMATRIX row-vector transform",
"implementation": "Godot Transform3D column-vector transform",
"invariant": "actor, offset, and bone composition produces the same world-space transform",
"tests": ["project/example_transform_parity_test.gd"]
}
]
}
```
An adaptation may change APIs or representation, never gameplay rules, branch outcomes, timing authority, or data authority. Missing equivalence proof keeps the contract at `MAPPED` or `PARTIAL`.
## Status meanings
- `UNMAPPED`: reference behavior has no located implementation.
- `PARTIAL`: known material behavior or branches are missing.
- `MAPPED`: both sides are located but not fully compared.
- `STATIC_VERIFIED`: complete material call-chain and implementation-equivalence comparison is documented.
- `TEST_VERIFIED`: static equivalence verification plus meaningful passing tests.
- `IN_PROGRESS`: currently being audited; do not use as a long-term resting state.
- `STALE`: reference, implementation, or test evidence changed after verification.
- `REGRESSION`: a previously passing behavior test now fails.
- `BLOCKED`: concrete missing input or dependency prevents progress.
- `EXCLUDED`: not ported by design; requires `exclusion_reason` and replacement verification.
## Verification gates
`STATIC_VERIFIED` and `TEST_VERIFIED` both require:
1. Non-empty reference and implementation file lists.
2. Existing detailed contract document.
3. Every implementation-equivalence field set to `VERIFIED` with supporting detail in the contract.
4. Every platform adaptation documented with its invariant and focused tests.
5. No unresolved material branch in `remaining`.
`TEST_VERIFIED` additionally requires at least one existing automated test and `last_test_result` equal to `PASS`.
When a verified entry becomes stale, retain its previous evidence and history. Do not delete the entry or recreate it under a new ID.
## History events
Append one compact JSON object per meaningful state change to `audit/history.jsonl`:
```json
{"time":"2026-09-19T00:00:00Z","event":"status_changed","id":"combat.local.normal_attack","from":"MAPPED","to":"TEST_VERIFIED","reason":"call chain audited and tests passed"}
```
Never store credentials, packet payload secrets, binary captures, or generated screenshots in the ledger.
@@ -0,0 +1,49 @@
# Project map
## Reference client
- Root: `../40250/Server Client TMP4/ClientVS22/source`
- High-level game and packet flow: `UserInterface`
- actors, motion, maps, flying objects: `GameLib`
- rendering, input, collision primitives: `EterLib`
- effects: `EffectLib`
- Granny integration: `EterGrnLib`
- audio: `MilesLib`
- terrain: `PRTerrainLib`
Use the Visual Studio project and active preprocessor flags to distinguish reachable product code from disabled, obsolete, and third-party code.
## Current client
- GDScript runtime and tests: `project/`
- native extension: `extension/`
- format readers: `formats/`
- shared/native libraries: `libgr2/`
- extracted assets and tables: `assets/`
- numerical reference oracle: `oracle/`
- host-side tools: `tools/`
## Existing evidence
- `docs/CLIENT-GAP.md`: historical broad gap analysis; migrate useful claims into contracts rather than trusting status text.
- `docs/CLIENT-PARITY-AUDIT-AND-FIX-GUIDE.md`: existing methodology and high-risk domains.
- `docs/CLIENT-40250-PORT.md`: porting context.
- `docs/PARITY-GAP.md`: visual/rendering gap notes.
- `oracle/run-diff-suite.sh`: Granny numerical comparison suite.
- `project/test_*_parity.gd` and `project/*_test.gd`: existing tests; inspect assertions before treating them as evidence.
## Typical verification commands
Run a narrow Godot test with:
```bash
godot --headless --path project --script project/test_name.gd
```
Some scripts expect the path relative to `project/` instead:
```bash
godot --headless --path project --script test_name.gd
```
Follow the convention already used by the selected test. Always finish code changes with `git diff --check`.
@@ -0,0 +1,264 @@
#!/usr/bin/env python3
"""Validate, refresh, and summarize the repository's 40250 parity ledger."""
from __future__ import annotations
import argparse
import hashlib
import json
import sys
from collections import Counter
from datetime import datetime, timezone
from pathlib import Path
STATUSES = {
"UNMAPPED", "PARTIAL", "MAPPED", "STATIC_VERIFIED", "TEST_VERIFIED",
"IN_PROGRESS", "STALE", "REGRESSION", "BLOCKED", "EXCLUDED",
}
PRIORITIES = {"P0", "P1", "P2", "P3"}
PENDING = {"UNMAPPED", "PARTIAL", "MAPPED", "IN_PROGRESS", "STALE", "REGRESSION", "BLOCKED"}
EQUIVALENCE_FIELDS = {
"preconditions",
"branch_structure",
"algorithms_formulas",
"state_transition_order",
"constants_units",
"timing_event_sources",
"resource_data_sources",
"protocol_side_effects",
"interruption_failure_cleanup",
}
def repo_root() -> Path:
here = Path(__file__).resolve()
for parent in here.parents:
if (parent / ".git").exists():
return parent
raise SystemExit("cannot locate repository root")
def manifest_path(root: Path) -> Path:
return root / "audit" / "manifest.json"
def load_manifest(root: Path) -> dict:
path = manifest_path(root)
if not path.exists():
raise SystemExit(f"missing audit ledger: {path}")
return json.loads(path.read_text(encoding="utf-8"))
def write_manifest(root: Path, data: dict) -> None:
manifest_path(root).write_text(
json.dumps(data, ensure_ascii=False, indent=2) + "\n", encoding="utf-8"
)
def append_history(root: Path, event: dict) -> None:
event = {"time": datetime.now(timezone.utc).isoformat(), **event}
with (root / "audit" / "history.jsonl").open("a", encoding="utf-8") as handle:
handle.write(json.dumps(event, ensure_ascii=False, separators=(",", ":")) + "\n")
def resolve_files(root: Path, data: dict, section: str) -> list[Path]:
contract = data
files = contract.get(section, {}).get("files", []) if section != "tests" else contract.get("evidence", {}).get("tests", [])
base = root
if section == "reference":
ref_root = Path(data.get("_reference_root", ""))
base = ref_root if ref_root.is_absolute() else root / ref_root
return [(base / item).resolve() for item in files]
def fingerprint(paths: list[Path]) -> str:
digest = hashlib.sha256()
for path in sorted(paths, key=lambda p: str(p)):
digest.update(str(path).encode())
if not path.is_file():
digest.update(b"<missing>")
continue
digest.update(hashlib.sha256(path.read_bytes()).digest())
return digest.hexdigest()
def validation_errors(root: Path, manifest: dict) -> list[str]:
errors: list[str] = []
seen: set[str] = set()
if manifest.get("schema_version") != 1:
errors.append("schema_version must be 1")
if not isinstance(manifest.get("contracts"), list):
return errors + ["contracts must be an array"]
for index, contract in enumerate(manifest["contracts"]):
label = contract.get("id", f"contracts[{index}]")
for field in ("id", "title", "subsystem", "priority", "status"):
if not contract.get(field):
errors.append(f"{label}: missing {field}")
if label in seen:
errors.append(f"{label}: duplicate id")
seen.add(label)
if contract.get("priority") not in PRIORITIES:
errors.append(f"{label}: invalid priority {contract.get('priority')}")
if contract.get("status") not in STATUSES:
errors.append(f"{label}: invalid status {contract.get('status')}")
if contract.get("status") == "EXCLUDED" and not contract.get("exclusion_reason"):
errors.append(f"{label}: EXCLUDED requires exclusion_reason")
if contract.get("status") in {"STATIC_VERIFIED", "TEST_VERIFIED"}:
reference = contract.get("reference", {}).get("files", [])
implementation = contract.get("implementation", {}).get("files", [])
evidence = contract.get("evidence", {})
if not reference or not implementation:
errors.append(f"{label}: verified status requires mapped files")
contract_path = evidence.get("contract", "")
if not contract_path or not (root / contract_path).is_file():
errors.append(f"{label}: verified status requires an existing contract document")
equivalence = contract.get("equivalence", {})
missing_equivalence = sorted(
field for field in EQUIVALENCE_FIELDS if equivalence.get(field) != "VERIFIED"
)
if missing_equivalence:
errors.append(
f"{label}: verified status requires VERIFIED implementation equivalence for "
+ ", ".join(missing_equivalence)
)
adaptations = contract.get("platform_adaptations", [])
if not isinstance(adaptations, list):
errors.append(f"{label}: platform_adaptations must be an array")
else:
for adaptation_index, adaptation in enumerate(adaptations):
missing = [
field for field in ("reference", "implementation", "invariant", "tests")
if not adaptation.get(field)
]
if missing:
errors.append(
f"{label}: platform_adaptations[{adaptation_index}] missing "
+ ", ".join(missing)
)
for test in adaptation.get("tests", []):
if not (root / test).is_file():
errors.append(
f"{label}: platform adaptation test does not exist: {test}"
)
if contract.get("remaining"):
errors.append(f"{label}: verified status cannot have remaining material branches")
if contract.get("status") == "TEST_VERIFIED":
evidence = contract.get("evidence", {})
tests = evidence.get("tests", [])
if not tests or any(not (root / test).is_file() for test in tests):
errors.append(f"{label}: TEST_VERIFIED requires existing tests")
if evidence.get("last_test_result") != "PASS":
errors.append(f"{label}: TEST_VERIFIED requires last_test_result PASS")
return errors
def command_validate(root: Path, manifest: dict, _args: argparse.Namespace) -> int:
errors = validation_errors(root, manifest)
if errors:
for error in errors:
print(f"ERROR: {error}")
return 1
print(f"PASS: audit ledger is valid ({len(manifest['contracts'])} contracts)")
return 0
def command_report(root: Path, manifest: dict, args: argparse.Namespace) -> int:
contracts = manifest["contracts"]
by_status = Counter(item["status"] for item in contracts)
by_priority = Counter(item["priority"] for item in contracts)
lines = ["# 40250 parity audit coverage", "", f"Total contracts: {len(contracts)}", "", "## Status", ""]
for status in sorted(STATUSES):
lines.append(f"- {status}: {by_status[status]}")
lines.extend(["", "## Priority", ""])
for priority in sorted(PRIORITIES):
lines.append(f"- {priority}: {by_priority[priority]}")
pending = sorted(
(item for item in contracts if item["status"] in PENDING),
key=lambda item: (item["priority"], item["status"], item["id"]),
)
lines.extend(["", "## Pending", ""])
lines.extend(f"- [{item['priority']}] {item['id']}: {item['status']}" for item in pending)
if not pending:
lines.append("- None")
report = "\n".join(lines) + "\n"
if args.write:
out = root / "audit" / "reports" / "coverage.md"
out.parent.mkdir(parents=True, exist_ok=True)
out.write_text(report, encoding="utf-8")
print(out)
else:
print(report, end="")
return 0
def command_next(_root: Path, manifest: dict, args: argparse.Namespace) -> int:
pending = sorted(
(item for item in manifest["contracts"] if item["status"] in PENDING),
key=lambda item: (item["priority"], item["status"], item["id"]),
)
for item in pending[: args.limit]:
print(f"{item['priority']}\t{item['status']}\t{item['id']}\t{item['title']}")
return 0
def command_refresh(root: Path, manifest: dict, args: argparse.Namespace) -> int:
changed = 0
ref_root = manifest.get("reference_root", "")
for contract in manifest["contracts"]:
working = {**contract, "_reference_root": ref_root}
current = {
"reference": fingerprint(resolve_files(root, working, "reference")),
"implementation": fingerprint(resolve_files(root, working, "implementation")),
"tests": fingerprint(resolve_files(root, working, "tests")),
}
old = contract.get("fingerprints", {})
contract_changed = old != current
reasons = [name for name, value in current.items() if old.get(name) and old[name] != value]
if reasons and contract.get("status") in {"STATIC_VERIFIED", "TEST_VERIFIED"}:
previous = contract["status"]
if args.write:
contract["status"] = "STALE"
contract["stale_reasons"] = reasons
append_history(root, {
"event": "status_changed", "id": contract["id"], "from": previous,
"to": "STALE", "reason": "fingerprint changed: " + ",".join(reasons),
})
if contract_changed:
changed += 1
if args.write:
contract["fingerprints"] = current
if args.write and changed:
write_manifest(root, manifest)
print(f"refresh: {changed} change(s){' written' if args.write else ' detected'}")
return 0
def main() -> int:
parser = argparse.ArgumentParser()
sub = parser.add_subparsers(dest="command", required=True)
sub.add_parser("validate")
report = sub.add_parser("report")
report.add_argument("--write", action="store_true")
next_parser = sub.add_parser("next")
next_parser.add_argument("--limit", type=int, default=20)
refresh = sub.add_parser("refresh")
refresh.add_argument("--write", action="store_true")
args = parser.parse_args()
root = repo_root()
manifest = load_manifest(root)
errors = validation_errors(root, manifest)
if errors and args.command != "validate":
for error in errors:
print(f"ERROR: {error}", file=sys.stderr)
return 1
return {
"validate": command_validate,
"report": command_report,
"next": command_next,
"refresh": command_refresh,
}[args.command](root, manifest, args)
if __name__ == "__main__":
raise SystemExit(main())