58-frame ir
session_20260902_17154658 tripod positions, raw IR stereo pair kept so FoundationStereo can run offline. The depth branch here is a network, not the camera.
| frames | 58 |
| capture mode | manual |
| stream | ir |
| projector | strobe |
| shot |
sceneforge pipeline
COLMAP poses, Gaussian-splat background, TSDF mesh, MuJoCo scene. The full nine stages.
lab58 all nine stages →
views 58pose backend colmapscale m/unit 0.2766verify PSNR 22.8gaussians 1,000,000scene compiles



















































































lab58_sam3d all nine stages →
views 58pose backend colmapscale m/unit 0.2766gaussians 1,000,000scene compiles









splat trainer ablations
Controlled Gaussian-training experiments on the same poses and held-out views; not production replacements.
30k MCMC background-only carve (diagnostic; no replacement assets)



30k MCMC background-only carve v2 (expanded silhouettes + metric depth)



30k full-res MCMC (0.01 opacity/scale; experimental)






30k full-res MCMC + gsplat antialiasing (control)

Depth-augmented PGSR 30k (evaluated; rejected)
held-out PSNR 26.7LPIPS 0.161median depth error mm 20.900p90 depth error mm 88.950free-space violation >50mm 10.5%planar gaussians 827,351
Rejected for the production pipeline: appearance is slightly better than corrected gsplat (26.74 vs 26.39 dB), but metric depth is worse (20.90 vs 18.01 mm median) and the novel-route flythrough has severe table/robot smearing and floaters. The run genuinely used metric FoundationStereo depth plus PGSR normal and multi-view consistency. Its 1 cm TSDF mesh is retained as research evidence, not adopted.



Depth-augmented RaDe-GS 30k (rejected)
held-out PSNR 22.5median depth error mm 24.330visible anisotropy >20 77.5%gaussians 817,535
Rejected: worse than corrected gsplat (26.39 dB, 18.01 mm median depth), and extreme anisotropy remains. Official Open3D mesh extraction segfaulted.


asset extraction candidates
Compact visual comparisons from the current scene capture; metric TSDF versus generative completion.
SAM 3D Objects (Meta) -- best generative asset on both objects
green cup: median vs measured depth 1.83 mm (RecGen 2.88, TRELLIS.2 5.79)small bottle: median vs measured depth 0.59 mm (TSDF 0.70, RecGen 3.72)green cup within 5 mm 69.5%small bottle within 5 mm 96.0%inference / peak VRAM 19-25 s, 22.8 GBlicence SAM License (not research-only)
Meta's SAM 3D Objects turns one masked image into a watertight textured mesh, a Gaussian splat and CoACD collision hulls in a single pass. Input is the stage-30 mask plus the best view from tools/rank_object_views.py, so it consumes what this pipeline already produces. Scored against the FoundationStereo depth measured at each object -- the only defensible question, since every generative method here hallucinates the unobserved side and none returns metres. It beats RecGen and TRELLIS.2 on both objects, and on the bottle it beats the measured TSDF mesh (0.59 vs 0.70 mm) by closing the surface where the sensor saw nothing; that rewards a plausible closed surface and does not verify the invented half. The generated scale is REFUSED and recorded in assets/report.json, exactly as tools/recgen_objects.py refuses RecGen's -- measured depth stays the metric reference. An earlier run put SAM 3D last on both objects; that was two adapter bugs (82.6 deg upright error, wrong axis proportions), fixed, and those numbers are void. The support/stool is still running and scored 9x worse than the measured mesh before the fixes -- unproven on large, half-occluded geometry. Full write-up in runs/sam3d_lab58/RESULT.md.









TRELLIS.2-4B current-data textured meshes (512; cup usable, bottle rejected)
method Microsoft TRELLIS.2-4Blicense MITcup surface retained 99.985%bottle trials 3 / 3 rejected
Official single-image TRELLIS.2 inference from a tight SAM RGBA crop. The turntable uses its native PBR surface renderer (not point samples). Detached faces were removed after welding UV seams; raw generation and depth-derived metric-scale GLBs are preserved. Cup view 000051 is recognizable with a rough rim. All three bottle views collapse to thin cards/slabs, so the bottle result is explicitly rejected. Visual screening only.






current-data assets: metric TSDF and RecGen









lab58 support/table asset: RGB-D extraction and simulation proxy
RGB-D vertices 41,029RGB-D faces 77,166registered texture views 58observed vertex coverage 100%
Two explicit representations of the captured wooden support. The raw metric RGB-D mesh preserves measured geometry but has ragged legs/feet. The clean metric box proxy uses the measured top/leg dimensions and an observed wood texture; it is the current MuJoCo visual/collision asset. A dedicated support capture is still needed for a polished free-standing mesh.


end-to-end demonstrations
Integrated scene, assets and deterministic MuJoCo task videos. Proxies and measured quantities are declared.
RoboSnap-style placement: ICP pose, support contact, decimation
ICP fitness (cup / bottle) 1.000 / 1.000ICP inlier RMSE 3.2 mm / 2.4 mmsupport top (fitted, 398k pts) 0.632 mcross-check: objects' measured cloud base 0.626 mstage-70 objects sank into table by 49-66 mmmesh budget 115 MB -> 4.8 MB, error 0.011-0.037 mmcup silhouette IoU 0.879
Four ideas ported from RoboSnap (vendor/robosnap) onto inputs it does not have. RoboSnap reconstructs from ONE photograph, so it invents the scene frame with VGGT and erases objects with a Gemini inpaint; both are skipped here because 58 registered stations, FoundationStereo depth and a trained splat are strictly better. What was taken is the part this pipeline lacked: until now stage 70 wrote a position and extents and NOTHING estimated object orientation, so a generated mesh landed pointing an arbitrary way. ICP against measured depth now registers each mesh (SE(3) only -- scale stays whatever depth said). Gravity comes from the fitted top slice of the support: 0.632 m from 398k measured points, cross-checked against the base of the objects' own measured clouds (0.626 m), since those objects physically rest on that surface. An earlier version of this entry claimed 6.2 mm agreement with the stool's 0.678 m height -- that was wrong. 0.678 m is the stool's *extent*, which only equals its top height if its base sits exactly at z=0, and it does not. The agreement was a coincidence. The projected-silhouette solve is a check, not a mover: given a leaking mask it slid the bottle 5.8 m, so it is capped at one object diameter. Decimation is not in RoboSnap and was written: 115 MB of generated meshes to 4.8 MB, still watertight, surface error 0.011-0.037 mm -- roughly 100x below the error the meshes already have against depth. Full detail in runs/sam3d_lab58/RESULT.md.
classical bottle pick v1: targeted RecGen 000006
finger contact 3.03 sverified lift 320 mmasset source targeted RecGen view 000006
Uses the stronger targeted bottle candidate (000006; 000041 has a white side artifact), metric in-scene XY and support-contact Z, and a stable box collider. Generic analytic proxy arm; weld-assisted after verified contact.


classical bottle pick v2: representative Franka Panda
finger contact 2.43 sverified lift 267 mmmax IK error 0.42 mm
Official MuJoCo Menagerie Franka Panda model used as a representative robot; captured robot CAD remains unavailable. Classical position-IK joint path. Kinematic attachment begins only after verified finger contact; a dynamic weld was rejected because it imparted an unphysical impulse.


classical bottle pick v3: image-matched UR5e + Robotiq 2F-85
finger contact 2.63 sverified lift 254 mmrobot UR5egripper Robotiq 2F-85static asset textured TRELLIS green cup
Official MuJoCo Menagerie UR5e and Robotiq 2F-85 assets replace the Franka placeholder. Base XY comes from the low-Z centroid of all registered robot masks with metric depth; the 0.55 m pedestal mount and 66.2-degree yaw are visually fitted to capture frame 38. The alignment sheet shows capture versus simulation at the same camera pose. The four-view video synchronizes registered external frames 21/38/51 with a side-mounted wrist camera. Coarse collision proxies are hidden (they caused the reported rotating joint-0 artifact), while Robotiq pads remain active. The metric TRELLIS cup is PCA-aligned upright and rests at its RGB-D tabletop position while the bottle moves. Classical IK drives the arm; bounded kinematic attachment begins only after verified Robotiq pad contact.








classical bottle pick v4: SAM 3D cup and bottle at measured poses
SAM 3D assets cup and bottle onlystool hand-fitted box, texture baked from the captureobject poses ICP against measured depthcup / bottle seating error 0.1 mm / 1.0 mmbottle mask-bbox IoU 0.00 (fit failed; xy is a fallback)cup mask-bbox IoU 0.676
Only the two manipulable objects are SAM 3D. The stool is still the hand-fitted box proxy with a texture baked from the capture, because SAM 3D's stool is the one asset it loses on: 33.09 mm median against measured depth versus 3.67 mm for the stage-60 mesh, 13.1% within 5 mm versus 58.3% (runs/sam3d_lab58/compare_support.json). Large, half-occluded furniture is where single-view generation stops paying. Same take as v3, with the hand-fitted TRELLIS cup and RecGen bottle replaced by the SAM 3D meshes at the poses tools/refine_placement.py measured. Three things were wrong before and are fixed here. The demo re-origins the cup mesh on export (PCA upright, xy centre, base at z=0) so a pose fitted to the generator's own frame put the cup metres off the stool; the export transform is now composed out of the measured pose. Each object is then seated on the box support proxy by its own lowest vertex, because the measurement rests on the reconstructed support plane and the two differ by ~6 mm. The bottle's cap proxy and collision box were sized and centred for the RecGen mesh and are now measured from whichever mesh is loaded. Honest caveat: the bottle's placement has a mask-bbox IoU of 0.00 -- that fit failed and its XY is a fallback, so its agreement with the photograph is luck, not measurement. The cup's is 0.6765. Second caveat, in the compositing rather than the placement: in views 21 and 51 the splat's own stool wins the depth key along the far edge and hides the bottom of the cup. The sim-only alignment sheet beside this one is the same scene with the splat switched off, and the cup stands clear on the wood there.




classical cup pick v1: feasibility pass
finger contact 3.13 sverified lift 290 mmtrajectory approach / close / lift / translate
Current-data feasibility demo. The arm/gripper is a generic analytic proxy, not the photographed robot; the grasp is weld-assisted after verified finger contact. Support scale/placement and cup scale/placement are metric.


classical cup pick v2: visual-quality pass
finger contact 3.13 sverified lift 290 mmtrajectory approach / close / lift / translate
Improved framing, arm placement/materials, observed-RGB support texture, and removal of duplicate simulated shadows. Generic analytic proxy arm; weld-assisted after verified finger contact. Metric cup and support placement.


same pick from six more capture stations (eye-level and overhead)
eye-level panes stations 3, 25, 8overhead panes stations 34, 53, 39eye-level ring elevation 8-12 deg, 0.72-0.98 moverhead ring elevation 22-38 deg, 0.62-0.96 mwrist camera unchanged in both
The same take, re-rendered from six stations the published 21/38/51 video never shows, three per video, with the eye-in-hand pane left alone so the grasp stays comparable across all three videos. The capture is two rings of tripod positions -- 0-31 at about 0.78 m and 32-57 at about 1.0 m -- so the two videos split along that: an eye-level trio and an overhead trio. Stations were chosen by how squarely each looks at the stool and how close it stands, then spread at least 90 degrees apart in azimuth from each other and from the published set. Nothing about the scene, the assets or the physics changes between the three videos; only which registered cameras the panes look through, which is the point -- every one of the 58 stations is a pose the splat and the simulation already share.






stool asset: SAM 3D vs measured mesh vs box proxy (same take)
SAM 3D stool, median vs measured depth 33.09 mm (measured mesh 3.67)SAM 3D seat height across the seat 0.579 - 0.633 m (54 mm bowl)SAM 3D seat under the cup / bottle 35.5 mm low / 4.0 mm highmeasured mesh seat under cup and bottle hole - never observedbox proxy seat flat 0.629 m
The same pick rendered with each candidate stool, because 'use the SAM 3D table' turns out to have a real cost. SAM 3D returns a complete stool but the wrong one: A-frame legs where the real stool has four square ones, and a seat that is a 54 mm bowl rather than a plane, so an object seated at the true tabletop height floats 35.5 mm over it at the cup's position and sinks 4.0 mm at the bottle's. The measured stage-60 mesh has the right silhouette and its own photographed wood, but a ray cast straight down at either object's position passes through it -- the stool surface under each object was occluded by that object during the whole capture, so there is nothing there to reconstruct, and the legs are reconstruction noise below the seat. The box proxy that v4 ships is flat at 0.629 m, which is where the measured mesh puts the seat everywhere it did see it (median 0.6262, p95 0.6294). That is why physics keeps the box in all three videos; only the visual changes. Filling that occluded patch is a job for the generative asset, and it is the one thing SAM 3D got wrong here.




VGGT-1B (feed-forward)
One forward pass over all views: poses and geometry at once, no COLMAP, no segmentation. Scale from the depth map.
VGGT-1B
views 58points 3,000,000scale m/unit 1.7820scale spread 1.4%
