System architecture
Separate capture, common pose data, recognition and game integration.
Layers and ownership
CAPTURE
POSE
MOTION
BRIDGE
GAME
| Layer | Actual boundary | Responsibility |
|---|---|---|
| Core | Platform-neutral primitives and monotonic clock | Time and common foundations |
| Pose | IPoseProvider, PoseFrame | Provider-neutral joint observations |
| Motion | Calibration, geometry, intents | Interpret motion as explicit events |
| Runner bridge | RunnerGameInputBridge | Game context and command acceptance |
| Native iOS | Objective-C++ C ABI | Camera ownership and latest-frame handoff |
The pose boundary
Apple Vision and MediaPipe SDK types stay behind the common pose model. Providers deliver frame identity, capture time, tracking state and joints in a PoseFrame. Native inference and Unity game decisions have separate responsibilities.
Distinct recognition paths
The legacy engine uses calibrated pose features and configured thresholds. V3 research uses aspect-correct geometry, local references, temporal evidence and episodes. Runner Demo Assist is a separate profile with observations distinct from V3. The current runtime constructs the session with gameplayV2: true, demoAssist: true.
Direction of data flow
The camera owner captures a frame. The provider delivers joints. The recognizer emits events. The bridge evaluates game context, time and duplicates before calling the game’s existing action methods.
Technology scope
Motion code uses C# and Unity. The native iOS boundary uses Objective-C++ and a C ABI. The Apple Vision path uses AVFoundation and Vision; the MediaPipe research path isolates SDK symbols in a separate framework. This website itself uses Next.js, React, TypeScript and MDX.