← Back
apromisedland

apromisedland/Tame_3D

Tame3D research framework: evidence-grounded multi-agent 3D reasoning with scene-level conformal uncertainty alignment

View on GitHub ↗
Stars
79
Forks
0
Watchers
79
Open issues
0
Contributors
1
Language
Python
License
Apache License 2.0
Default branch
master
Created Sep 28, 2026Updated Sep 28, 2026

Star growth

Today—
This week—
This month—

Star history will appear here once this repo has been tracked for a couple of days.

README

Tame3D

Evidence-grounded multi-agent 3D situated reasoning with temperature alignment, scene-block conformal prediction, bounded acquisition, and abstention.

Research implementation of Taming Multi-agent Collaboration for 3D Situated Reasoning via Uncertainty Alignment. A manager plans typed geometric programs, an egocentric expert reads local views, a global expert reads geometry and a bird's-eye view, and a critic checks evidence before calibrated fusion.

Evidence status: this repository contains an executable foundation-model interface and offline validation. It does not contain measured SQA3D foundation-model results. The preserved five-seed study is a synthetic geometry simulation. The source presentation's “6%” improvement statement is retained as an unverified source claim, not as a result of this code; its baseline, split, predictions, and definition of percent versus percentage points are unavailable.

Quick start

Python 3.11 or 3.12:

python -m venv .venv
.venv/bin/python -m pip install -e ".[dev,scannet]"
.venv/bin/python -m tame3d demo --output runs/demo
.venv/bin/python -m pytest -q

Windows PowerShell:

& 'C:\path\to\verified\python.exe' -m venv .venv
& .\.venv\Scripts\python.exe -m pip install -e ".[dev,scannet]"
& .\.venv\Scripts\python.exe -m tame3d demo --output runs/demo
& .\.venv\Scripts\python.exe -m pytest -q

Use the project interpreter by path. Do not install project packages in a shared application runtime. The lightweight core needs no GPU, PyTorch, model weights, API key, or dataset download.

The demo writes scene-disjoint data, complete development/calibration trajectories, fitted temperatures, scene thresholds, evaluation predictions, metrics, response replay records, and rendered PNGs. Its geometric integration results and its separately labelled constructed stopping-rule fixtures have different purposes. The latter demonstrate immediate answering, acquisition followed by answering, and abstention; they are not model measurements.

中文快速入门 · Method and interfaces · SQA3D and models · Evidence and reproduction

Commands

Command Input and output
demo Complete offline integration exercise and constructed controller examples
prepare-sqa3d Official annotations plus ScanNet/prepared scenes → scene-disjoint JSONL datasets
collect Questions, scene store, frozen policy → complete potential trajectories
fit Development trajectories and separate labels → expert temperatures
calibrate Calibration trajectories, labels, temperatures → scene thresholds
infer Unlabelled questions, scenes, matching calibration → predictions and evidence
evaluate Predictions and held-out labels → metrics, scene-bootstrap intervals, optional paired differences
reproduce-synthetic Original five-seed simulation, independent audit, historical comparison

Run python -m tame3d COMMAND --help for arguments. configs/offline.yaml is offline. configs/http.yaml contains placeholder model identifiers: replace them with deployed text/vision model snapshots. Hosted endpoints can specify api_key_env; only the variable's name is stored.

Research behavior

  • Tools use metres and an observer frame with +x right, +y forward, +z up. SQA3D center offset, axis alignment, quaternion order, and base forward axis are converted on input.
  • Egocentric tools access only objects visible in accumulated views. Global tools access the object table and metric geometry. Oracle instances and external predictions are distinct regimes.
  • Programs are typed operation graphs, never executable model-generated Python. Primitives are detect, select, transform, visible, relate, count, and nearest.
  • Experts score shared per-round candidates plus <OTHER> using fixed-count sampled choices and smoothing. Failed experts receive a fixed uniform vector.
  • Development-only temperatures precede equal-weight fusion. Calibration takes each scene's worst question at each potential round.
  • Only an ordinary singleton is accepted. Residual candidates, ambiguity, and empty sets trigger acquisition or abstention. Forced answers are separate.
  • Configuration, prompts, implementation, model identifiers, pose/instance regimes, and acquisition parameters are fingerprinted. Mismatching artifacts are rejected.

Under exchangeability of whole scene/question blocks and a frozen policy, the construction bounds the probability of any wrong accepted answer in a new scene. It does not bound error conditional on answering, certify arbitrary shifts, or establish a real-scene accuracy improvement.

Historical experiment

python -m tame3d reproduce-synthetic --output runs/synthetic
python -m tame3d reproduce-synthetic --output runs/synthetic --verify-only

The full run uses seeds 17, 29, 43, 71, and 101, with 400 development, 600 calibration, 1,000 clean test, and 1,000 shifted test scenes per seed. This is separate from the small integration demo. Historical summaries are included; generated trajectories remain local. NumPy 2.3.5 is the historical numerical dependency.

CI runs tests, an offline demonstration, and package builds on Windows/Linux with Python 3.11/3.12. The full study is an explicit reproduction command.

Boundaries

The global backend is a VLM over a bird's-eye image and geometry table, not ShapeLLM or a native point-cloud model. Detection uses supplied instances; no detector, segmenter, or language-based localizer is trained here. External predicted poses require their own calibration.

Point-cloud z-buffer visibility is a projection approximation, not dense mesh visibility. Scene-maximum calibration can be conservative. Remote models may not honor seeds exactly; save immutable snapshots and replay records and recalibrate when behavior changes.

Apache-2.0. Third-party notices identify the pinned SQA3D normalizer. Original manuscripts, slides, restricted datasets, weights, and credentials are excluded.