Exam surface
Author, assemble, run, and replay timed stations for learners, faculty, and admins— with traces, actor turns, and review packets in the loop.
Clinical skills simulation · WebXR-first
Timed clinical skills stations—from case blueprint to WebXR runtime— with faculty review, durable traces, and promotion bound to the reviewed revision.
Inspired by Step 2 CS-style multi-station flow. Built as an encounter factory, not a pile of one-off scenes. Not an exam-equivalence or clinical-validity product.
Authors define stations. Learners run timed encounters. Faculty review traces and packets. Admins hold the gates.
Authored content identity now follows an encounter through faculty promotion, a pinned runtime bundle, timed learner phases, actor execution, and review.
Promotion and readiness flags stay false until hardware and policy say otherwise. We document limits in public.
Cleared movement evidence · 2026-09-04
This six-second browser capture shows skeletal motion reaching the head, torso, shoulders, arms, and hands. It is a rig-deformation proof, not a polished performance.
The visible body, hair, shirt, and motion sources cleared the project’s publication review. The lower-body crop is deliberate: the source trousers have an unresolved per-item license conflict, so that geometry is not published here. Viseme footage is also withheld until its separate source provenance is resolved.
Machine-readable provenance · license review
notEvidenceFor: production animation, cloth behavior, speech or viseme synchronization, Quest readiness, clinical validity, or exam equivalence.
Speech · emotion · hair · skin · 2026-09-19
Thirty-four seconds of the factory scratch nurse: three spoken lines, an authored emotion arc, blinking, strand-textured hair, and pore-level skin. The file includes an audio track. Turn the sound on. CEO grade: sil sealed; PP sealed (first time); aa open with upper tooth row; lower cavity still a dark hole.
Progress log entry 109 · scratch bake, not the shipped GLB
notEvidenceFor: production TTS, Quest readiness, clinical validity, exam equivalence, gums/cavity detail, the shipped GLB, or a live station-reply capture.
Platform
Author, assemble, run, and replay timed stations for learners, faculty, and admins— with traces, actor turns, and review packets in the loop.
Reviewed case definitions drive scenes, actors, dialogue posture, emotion timelines, asset needs, and persistence—so work scales past hand-built demos.
Rooms, humanoids, clothing, equipment, and provenance intended for reuse across encounters— with cage-match style comparison before anything is treated as “ready.”
Sidecars for humanoid generation, voice, IWSDK spikes, and providers stay isolated until evidence supports a promotion decision. Experiments don’t become product by accident.
Current state · 10 September 2026
Source of truth for status is the repo ledger (PROJECT_STATUS.md), not this page.
Public copy here is a snapshot for humans—updated when the product story actually moves.
Measured, not asserted · 2026-09-10
Nine cards landed against a frozen acceptance contract. Each carries an evidence report graded by its own verifier against retained artifacts, and each was re-measured by the owner against the tree rather than accepted from its report. Three of nine were rejected at least once; one was rejected three times. The rejections and the failures are published alongside the passes, because a report that only lists passes is not evidence.
| Metric | Measured | Threshold | Outcome |
|---|---|---|---|
| Arrival error | 0.03597 m | ≤ 0.05 m | pass |
| Settled heading error | 0.000° | ≤ 10° | pass |
| Stopped duration | 4.816 s | ≥ 2 s | pass |
| Root travel while stopped | 0.000000 m | 0 m | pass |
| Deepest floor penetration | 0.00179 m | ≤ 0.005 m | pass |
| Foot slide | 8.00 Hz capture | ~132 Hz needed | not gradeable |
The last row is the useful one. Identifying a foot-contact window at the rubric's 0.005 m per-frame allowance needs roughly 132 Hz; four skinned humanoids and a compiled room through software rendering gave 8.00 Hz. Rather than grade an aliased stream, the instrument refuses and reports the number. An earlier revision of that report graded it as failing from an assumed 60 Hz that was never measured; that verdict was wrong in the same way a passing one would have been, and is retracted in place in the published report.
Other findings worth naming: the shipped locomotion clip fails the frozen rubric, and no clip this pipeline has produced meets the foot-plant threshold. A learned-motion candidate was screened and held, not adopted — its text encoder depends transitively on a gated model licence this project does not hold, and the host has no CUDA device. Twenty-eight open items are declared across the nine reports.
Reports: SC-00 rubric · SC-01 · SC-01S · SC-02 · SC-03 · SC-04 rights · SC-05 approach · SC-06 replay · SC-10 research hold
notEvidenceFor: clinical validity, scoring validity, exam equivalence, and worn-headset readiness are not claimed and have no evidence here. Three of twelve acceptance rows — an uninterrupted workflow recording, a website demonstration, and independent final acceptance — are unstarted, so this package is not closed. No humanoid imagery from these cards is published: their public render rights are blocked pending a named upstream licence resolution.
Evidence you can look at · 2026-08-31
These screenshots were captured from commit 05120405, where
BothyBoard children W1–W19 (plus W14a/W14b/W11s) had landed.
Faculty author the encounter as a Scenario, then review a World Compile Graph
of baker families. Add node / Remove node mutate the graph; the lock table stays
the lock write path. Chromium Playwright of the running admin app
(ui-admin + local API), not a schematic.
Since this capture, authored revision identity now invalidates stale approval,
and Admin can preview the runtime delta before a reviewed revision is promoted
into an immutable encounter bundle.
claimScope: local admin UI capture of an authored Scenario at
capture commit 05120405 plus the compile DAG recaptured 2026-09-10
at bbf3f84a; current revision identity and runtime-delta behavior
are supported by later committed application tests, not by these stills.
notEvidenceFor: Quest readiness, clinical validity, live Blender bake,
LLM draft quality, or that the parent BothyBoard program card is closed (it is still Idle).
Node labels truncate in this crop; that is the canvas, not missing data.
Evidence you can look at · 2026-09-01
Each production factory station is a Standard Schema V1 interface
(~standard.validate plus jsonSchema.input).
Admin cards are derived from that schema: one card per id, one control per
property, Apply refuses invalid payloads.
instrument is a gate and is not a card.
The TRELLIS control adds a worldview bake model (subject/pack
ecg-cart-imagine-box), distinct from a fixtureSlot bind.
FactoryStationCards
(Vite 127.0.0.1:5174/factory-station-cards-shot.html), 688×3870 native,
captured 2026-09-10 at HEAD bbf3f84a.
Subject: all nine production station cards. Scroll the figure for the full stack.
claimScope: local admin React cards derived from
factoryStationSchemas on origin/main.
notEvidenceFor: live TRELLIS GPU bake, Quest readiness,
clinical validity, or that Apply invokes a baker (it validates and emits
the payload).
Evidence you can look at
Three of the fourteen stations lay their patient down. For weeks those three rendered a crumpled knot on the mattress. Five separate fixes landed and none of them changed a pixel — each was aimed at a cause nobody had located. On 2026-08-20 the sixth worked, because an ablation was run first instead of a sixth guess: render the pose's two mechanisms separately and see which one is the damage.
The cause was seventeen joint rotations hand-tuned against Anny's 23-bone skeleton, being applied to an MPFB body with 138 joints. They are correct for the rail they were written for, so the fix was not to delete them — it was to scope them. Anny keeps its table; MPFB skips it. One switch, and the right-hand cell above became the centre cell.
The part worth keeping is the measurement that could not tell them apart. Before the fix, the knot and the person had identical envelope metrics — height 446 mm, minimum Y 0.570, to three decimal places. One was a patient and one was refuse, and every bounding-box assertion in the repository scored them the same. Any future contract that grades a lying figure by its bounding box is measuring nothing.
Not fixed, and visible in that capture: her arms rest raised in the air rather than at her sides. That is the bind pose showing through now that the wrong rotations are gone, and the answer to it is a retargeted recumbent motion clip — not a second table of hand-authored angles. That work is open. Nothing here is a claim about clinical validity of the pose; it is a claim that a figure on a bed reads as a human being.
Evidence you can look at
A case definition describes a person in numbers: age, height, build, how tense they look. Until today the factory did not read them. Bodies were baked from hand-written Python presets — four of them, covering four actors out of a hundred and four. Every other figure a learner met was the same body with a different job title, and the numbers in the case were decoration.
The case now authors the phenotype and the generator resolves it. Measured in the shipped fixture data: 32 of 32 actor records carry numeric phenotype, where before the audit found thirty-eight of forty-two actors with none at all.
What this is not: it is not a claim that every shipped body has been re-baked from these numbers. It is the data and the seam that reads it — the resolver now has something to resolve. Re-baking the cast against it is the next piece of work and it is open. Nothing here is a claim about clinical validity of any figure; it is a claim that the case definition finally reaches the body.
Evidence you can look at
A case definition describes a person: age, build, hair colour, eye colour. The factory was
reading almost none of it. Every shipped figure's iris came from a role default —
patients brown, family green, nurses blue — assigned by matching words in a job title. The
case's own eye_color field was carried all the way to the bake and then
dropped, because the one call site passed an empty dictionary where the phenotype belonged.
The more useful half is what happens when a case asks for something the factory cannot build. One patient is authored with hazel eyes. There is no hazel material in the licensed set, and the old behaviour was to silently fall back to a role default — the case asked for one thing, the learner saw another, and nothing anywhere said so. The factory now refuses that value out loud, and publishes the nine colours it can actually produce, each with its licence, generated from the same list the renderer reads so the two cannot drift apart.
That refusal is deliberate and it is the point. A factory that quietly substitutes is a factory whose output nobody can trust to match the case; a factory that refuses forces the question upstream, where a human or an authoring tool can answer it against a real list of options. The hazel patient is still unbuilt, and that is the correct state until someone chooses.
Not claimed: that these faces look finished. The skin is procedurally baked and carries no painted facial detail, which is visible in these crops and tracked separately. This shows one field of the case definition reaching one property of the render, verified by comparing the shipped bytes and by looking at the pixels.
Evidence you can look at · 2026-08-28
Shipped MPFB bodies carry MakeClothes library garments fitted onto the standard rig:
teal scrubs on the nurse, a cyan exam gown on the adult patient, sage street clothes
and boots on the walk-in, muted rose on the family partner. Contrast against skin is
visible at a glance. These are isolated grade captures of the tracked GLBs under
apps/ui-xr/public/generated-humanoids/, not the 2026-08-21 sage t-shirt
pair that documented a palette bug.
Street still: isolated EEVEE finished_figure_grade.py --frame full with
world Background strength 0.55 (grade-lighting.json), 2026-09-10. Other cells:
model-vetting-glb-grade-capture front_lit, 2026-08-25/28. Garments are
fitted via MakeClothes ClothesService.fit_clothes_to_human onto the MPFB
basemesh. A 2026-08-21 incident (cream closed_casual tint 21 RGB units
from skin) is closed.
Still visible and not claimed as production cloth: nurse and family
still show a waistband gap (cargo cover shell). Street trousers are MakeClothes
pants02 punkduck_male_classic_jeans (pack-page CC-BY), not a cover shell.
Isolated EEVEE bake includes world Background strength 0.55. Poke-through at the
chest on some shirts. Isolated harness, not a station.
claimScope: role-distinct fitted clothing on shipped MPFB
GLBs. notEvidenceFor: production cloth, Quest readiness, pregnancy
morphology, clinical validity, or exam equivalence.
Evidence you can look at
Every garment in the audited legacy MakeHuman/MakeClothes cache was fitted, rendered,
and assigned a geometry-derived class. Within that specific cache,
hospital_gown has zero members. This does not describe the
separate case-linked real-gown path shown elsewhere in the repository and on this page.
street 8 · footwear 3 · labcoat 2 · scrub 2 · evening_dress 1 ·
hospital_gown 0. Nothing fell into other and nothing was left unknown.
The first cell is the whole reason this was built. A file named
crudegown.mhclo once passed three machine contracts — its licence was verified,
its vertex indices were verified, its presence on the body was verified — and then somebody
looked at the render and found a floor-length spaghetti-strap evening dress. Presence,
placement and provenance are three questions, and none of them is class. The
inventory now answers the fourth one from geometry, in machine-readable form, before anyone
opens an image: crudegown → evening_dress, hem at 3% of body height.
The practical consequence is narrower: this legacy cache cannot satisfy a hospital-gown request, and the resolver must not substitute its evening dress because the filename sounds plausible. The current real-gown path is a separate, provenance-tracked source and still has to pass its own fit, motion, and runtime evidence gates.
Evidence you can look at
Garment layers now come from the case definition and are visible on browser-rendered actors, including the real-gown path and clothed multi-role stations shown on this page. That is evidence of case-to-runtime wiring, not a claim of production-quality cloth, fit, or motion.
Correction, 2026-08-10 evening. This section previously said one actor “renders translucent”. That was our diagnosis and measurement disproved it. Nothing on that figure is transparent; the garment covers every face of the region it claims — across four body/slot pairs, hidden-plus-behind-cloth accounts for all of them and zero faces have no garment nearby. The skin a viewer notices sits outside the claimed region, which makes it a question about how far the garment should reach, not a rendering fault.
The open quality bar is visible: some garments still read as rigid shells, hems and joins remain rough, and motion evidence does not yet establish production cloth behavior. New clothing claims therefore require measured Model Vetting and UI-XR evidence rather than inference from asset presence alone.
Inspection packet: ed-real-garment-webxr-inspection.json. Deeper factory reports live under docs/openclinxr and Model Vetting cagematch outputs in the repo.
Evidence you can look at
On 2026-08-11 a rig upgrade landed on the library bodies — MPFB's 64-bone
mixamo_unity skeleton and its shipped CC0 weight map, replacing a hand-rolled
bounding-box armature whose hands carried 0.00% of the skin weight with
20,216 vertices collapsed onto a single upper-arm bone. Three machine contracts passed.
Then the capture was graded by eye, and the figure had no trousers.
The revert took one commit. What it exposed took the rest of the day: the trousers were
never rebuildable. Scrub_Shirt.mhclo and cortu_cargo_pants.mhclo
were on no disk in the repository — the staging directory is empty and
gitignored, and the provider cache held three upper garments and no lower-body source at
all. The 2,530 triangles in the tracked asset came from a bake whose input no longer
existed. Meanwhile the pipeline's own “find-or-stop” guard, asked for a lower garment and
finding none, logged a warning and baked the body anyway.
Both halves are fixed and in the repository: the guard now throws instead
of warning, and the sources are acquired, tracked, and recorded in the licence ledger at
acquisition time — makehuman-pants01 (CC0, Cortu Johnstone)
and Scrub_Shirt (CC-BY, WojackOWL). A re-bake now finds its
inputs with no network, or refuses loudly.
What this picture is not. It is not a claim that the figure looks good. The same capture carries three defects we have measured and not fixed: a sawtooth band of bare skin where the shirt hem meets the trousers; hands painted the garment colour rather than clothed — the shirt mesh spans Y 0.913–1.505 m on a 1.760 m body while the hands sit near 0.79 m, below its lowest vertex; and a shirt with no volume of its own, breast and navel anatomy reading straight through it.
A correction we made to ourselves. Grading the lit pass alone, we recorded the hands as “mittens with no separated fingers” and were about to file missing hand geometry. The structure pass shows fingers and thumbs present and well formed. Lit resolves silhouette, structure resolves topology, and the lit pass flatters. The wrong finding was caught because both passes are captured, not because anyone was careful.
Captures produced by model-vetting-glb-grade-capture, whose NodeIO-versus-scene-graph
self-check agreed to 1.2 × 10⁻⁵ relative error. That agreement proves the renderer drew
the file and nothing about whether the file is right — which is why a human graded the pixels,
and why the missing trousers were noticed at all.
Evidence you can look at
The two library bodies were bound to a hand-rolled bounding-box armature with Blender's
automatic weights. Measured, that put 0.00% of the skin weight on
hand.L/R and collapsed 20,216 vertices onto a single upper-arm
bone — an arm that moved as one rigid piece from shoulder to fingertip. The fingers
were in the mesh. Nothing could move them.
They now ride MPFB's 64-bone mixamo_unity rig with its shipped
CC0 weight map — mixamorig:LeftHand alone carries 592 vertex
entries, with full finger chains. Both files shipped with MPFB and neither was being used.
Why supine, and why that is the better evidence. A standing rest pose looks identical before and after a skeleton change — it would prove nothing. These are recumbent renders from the isolated posture harness, and the folded arms and separated fingers are shapes the previous rig could not produce at all.
What these are not. Isolated lab renders, not in-station frames: no room, no other actors, no clinical context. Framing is computed from the subject's bounding box by the harness, not authored — which is the point, because the authored per-mode cameras in the full scene are a separate and still-open defect. Two attempts at hand-tuning those camera positions today were measured and reverted, one of them producing a 7.4 KB blank frame.
Defects still visible and unfixed: the sawtooth seam where the top meets the trousers, and low-polygon faceting across the limbs. Neither is claimed as solved.
Dark software factory · shaping up
Lights-out generation: a hard-surface prompt becomes a budget mesh through
named, measured stations — not a person pushing sliders on one GLB.
Grok Imagine packs condition TRELLIS Metal in an isolated process; then
factory:trellis:optimize runs high-error meshopt targets from the
raw mesh, a weld pass, and optional factory:trellis:pack (gltfpack).
One ECG cart, VR hard-surface pack, measured 2026-08-11:
973,639 raw → 60,000 preferred (clears ≤80k);
the same ladder can stretch to 34,443 (−96.5%) when a station
needs the share band. Photoreal packs on the same ladder stalled at ~186k.
Meta Quest 3 class scene guidance is ~1.3–1.8M tris for a whole scene;
≤40k is a multi-prop share, not the device limit. Not worn-headset readiness.
Not clinical realism.
Low-poly game-ready medical ECG monitor cart prop for WebXR / Quest. Hard-surface stylized, NOT photoreal. Clean boxy forms only. Wheeled base · upright column · large matte black screen. 6–8 square button pads · ≤6 circular jacks. NO free cables. NO logos, labels, text. Matte grey plastic. Studio grey bg. Maximize large flat planes for 3D reconstruction.
Hard-surface pack prompt — not a photoreal product shot. Photoreal inputs defeat post-opt (measured ~186k floor); hard-surface packs unlock preferred / share bands without fighting high-frequency detail.
factory:trellis:bake
MADR 0050: do not reject the generator on raw tris. Judge after optimization. Raw megameshes are never delivery assets.
Direct high-error targets from raw. Chain ratios plateau ~59k on this pack.
Stop at the first graded rung under ≤80k; 34k is share-band stretch, not “Quest wants 40k.”
Delivery: optional factory:trellis:pack. Harness prop — not a worn-headset claim.
| Station | What it does | Measured |
|---|---|---|
| Hard-surface pack | Grok Imagine prompt for boxy planes, not a photoreal product shot | Photoreal pack post-opt floor ~186k. Hard-surface chain ~59k; high-error stretch 34k |
| Isolated Metal bake | One OS process per subject (factory:trellis:bake) |
Same-process multi-subject TRELLIS OOMs the MPS heap and cascades |
| Multi-view condition | front / side / ±¾ PNGs concatenated into TRELLIS embeddings | ECG far-side fill 0.35 → 0.44, surface area +47%, 3.7% fewer triangles |
| High-error from raw | factory:trellis:optimize direct targets (180k / 120k / 80k / 60k / 40k) |
973,639 → 179,999 / 60,000 / 39,999. Chain ratios plateau; further 0.05 cuts barely move |
| Weld, then same targets | Position merge so split verts stop inflating the count | 40k rung stays 39,999 after weld; 25k stretch floors ~34.5k |
| Pack | factory:trellis:pack gltfpack quantize + meshopt compress |
Delivery size. Do not use -sa -se 1 (can zero the mesh) |
| Champion policy | First graded rung under preferred ≤80k; denser sibling if ≤40k looks worse | 60k is a legitimate champion. 34,443 is a measured stretch, not a Quest-3 prop limit |
Budget policy: prop preferred ≤80k (default stop) ·
share ≤40k only under multi-prop pressure · acceptable ≤120k · skeleton hard ≤180k
(partial station, not a full multi-actor exam) · Quest 3 device-class ceiling ~1.3–1.8M
scene tris (Meta native). Draw calls, materials, and fill-rate usually beat
another −15k on one cart. Order: batching → materials → KTX2 → multiview → tris.
CLIs: pnpm factory:trellis:bake →
pnpm factory:trellis:optimize →
pnpm factory:trellis:pack
(hatch: pnpm factory:trellis:hatch).
Skill: .agents/skills/trellis-vr-equipment-optimize/.
Evidence: .openclinxr/evidence/trellis-bake-vr-hard/,
trellis-vr-optimize-iterations/ecg-cart-vr-hard/.
claimScope: factory automation path for equipment props, measured on the ECG cart hard-surface pack.
notEvidenceFor: Quest readiness, clinical accuracy, exam equivalence, sampler-knob sweeps
(those fifteen knobs are still vendor balanced-tier defaults).
Evidence you can look at
The input half, added 2026-08-10. Before anything is reconstructed, the factory renders reference views of the object from its own parametric builder. On 2026-08-10 that step went from 3 subjects to 38: 35 clinical objects, five views each, 175 renders in a single 61-second pass with one browser and one dev server. Every view is measured for isolation — no room geometry, no HUD, no ground plane — and every one passed.
Each of these began as one reference image and was reconstructed to a textured mesh by TRELLIS running on Apple Silicon Metal — no hand-modelling, no per-asset artist pass. Rendered here in isolation from the shipped GLB, lit and wireframe, at the triangle budget each survives.
The wall clock is now consumed by the runtime: a station that declares it renders the generated mesh, measured at 34,885 triangles in a live scene, up from a 26-triangle placeholder. The monitor (~106k) and denser ECG cart variants remain isolated harness renders pending grade + multi-prop station share — 106k is still under prop acceptable (≤120k) and far under Quest 3 scene class (~1.3–1.8M). None of this is a Quest performance claim, and none of it is evidence of clinical realism.
Evidence you can look at · 2026-08-14
Last-resort factory when the object is not in the kit/parametric store and
not acquirable CC0/CC-BY. One Grok Imagine upper-¾ on a black void,
flood-keyed (Imagine writes JPEG RGB — no native alpha), then TRELLIS.2 Metal with
Space-order remesh on compact extracts, then factory:trellis:optimize to the
preferred ≤80k stop. Six subjects ran that path. The 49M medication cart is a recorded miss.
Kit remains exam SSOT. These GLBs are harness champions — not promoted into
ui-xr.
Orchestrator pixel grade, 2026-08-14: four Imagine plates are isolated black-void product shots;
four 80k champions are the same silhouettes with remesh-soft surfaces.
Pulse-ox and glucometer also ran the hatch to 80k; their grade crops are either too tight (hinge cavity)
or flank-on, so they stay in the ledger and off this page.
claimScope: escape-hatch factory stills for compact hard-surface props.
notEvidenceFor: Quest readiness, clinical accuracy, device equivalence,
kit replacement, UI-XR promote, Imagine-shader smoothness.
CLI: pnpm factory:trellis:hatch (commit c758a276).
Evidence: .openclinxr/evidence/trellis-escape-hatch/ (gitignored champions).
Next proof
The next work is to turn the newly connected lifecycle into visible, measured product evidence. Dates are not SLAs, and nothing here promises production deployment, headset certification, exam equivalence, or clinical validation.
Public roadmap copy follows committed product evidence and the repository’s protected claim boundaries.
The operational queue remains in PROJECT_STATUS.md and the project board.
Deployment posture
Development and validation run locally with deterministic providers and preconfigured assets where possible. Connected adapters exist behind explicit gates. Azure is a long-term home for API, orchestration, and admin surfaces— not a claim that production is live today.
Build model
The repo runs an OpenClaw-style operating model: role charters, path scopes, leases, drift guards, and slice records. That keeps long autonomous work from inventing features outside the blueprint-factory mission. It is how we build OpenClinXR—not a separate product you install.
Evidence Docs
Marketing pages should not replace the ledger. The following links are committed factory and evidence posts for operators and auditors.