Real-time open-vocabulary relation prediction from any inputs.
Paper · Project page · Models · Training data · Documentation
Give RelateAnything an image, object regions, and the relations you want to
look for. It returns scored (subject, relation, object) triplets, such as
person → riding → horse.
- Use your own regions. Boxes can come from a detector, a segmenter, or a person. Masks are optional.
- Choose the relation vocabulary. Supply phrases at inference time without retraining. A small text encoder embeds them once; no language model runs per frame.
- Get one graph or two. Return a single ranked list, or separate spatial and semantic graphs from the same forward pass.
hero_30s_five_clips.mp4
The same relation model across five scenes, with tracked regions and relations. Clip credits · Reproduce the video
Try the browser demo — no installation; inference runs on your device using WebGPU or WASM.
Quickstart · Models · Deployment · Performance · Results · How it works
Use Python 3.12+. Install a CUDA-enabled PyTorch build for NVIDIA GPU
inference, or use device="cpu" in the example below.
git clone https://github.com/Maelic/RelateAnything.git
cd RelateAnything
pip install -e ".[hub]"Run this from the repository root. The sample photo is included; the two example boxes identify the person and the horse.
import numpy as np
from PIL import Image
from relsgg import RelateAnything
model = RelateAnything.from_pretrained("maelic/relsgg-vits16plus", device="cuda")
image = Image.open("assets/reel/images/horse.jpg").convert("RGB")
boxes = np.array([[470, 130, 650, 630], [90, 310, 1010, 875]], dtype=np.float32)
# Boxes are [x1, y1, x2, y2] in original-image pixels.
# Labels are optional and only make the printed triplets easier to read.
for triplet in model.predict(image, boxes, box_labels=["person", "horse"], topk=5):
print(triplet)
# Change what relations to look for, without retraining.
model.set_vocabulary(["riding", "carrying a rider", "beside", "in front of"])
graphs = model.predict(image, boxes, box_labels=["person", "horse"], decompose=True)
print(graphs["spatial"])
print(graphs["semantic"])Released checkpoints include the backbone configuration and text encoder; no gated DINOv3 login is needed for inference. Replace the sample image and boxes with your own inputs; the API also supports masks and a larger built-in relation vocabulary.
API guide: inputs, vocabularies, masks, and scores · Installation and offline use
Start with relsgg-vits16plus, the recommended balance of quality and speed.
All three checkpoints use the same training recipe; their vision backbones differ.
| Checkpoint | Backbone | Parameters |
|---|---|---|
relsgg-vits16 |
DINOv3 ViT-S/16 | 46.1 M |
relsgg-vits16plus |
DINOv3 ViT-S/16+ | 53.2 M |
relsgg-vitb16 |
DINOv3 ViT-B/16 | 113.8 M |
Compare model quality and speed. Each model card includes evaluation results and calibration details.
- PyTorch: use the Python API above on CPU or CUDA.
- NVIDIA GPU / TensorRT 10: build a local FP32 engine from the released ONNX graph. Vocabulary selection, calibration and two-graph decoding stay available. Setup and examples.
- Laptop / ONNX or OpenVINO: run the local image or webcam demo. Setup · OpenVINO.
- Browser: try the demo without installing Python.
After downloading the relation graph and building the local detector as explained in the deployment guide:
# CPU image demo; omit --image to use a webcam.
python deploy/demo_webcam.py --dist deploy/dist/relsgg-vits16plus --image photo.jpg
# NVIDIA GPU, after building both TensorRT engines.
python deploy/demo_webcam.py --dist deploy/dist/relsgg-vits16plus \
--backend tensorrt --device cuda --image photo.jpg --decomposeThe relation model consumes regions; it does not detect objects itself. Demo detectors are rebuilt locally and have separate licenses. The exported ONNX/TensorRT graph scores boxes; native mask inputs are supported by the PyTorch API. Deployment contracts and limitations.
On an RTX 3080 Laptop GPU, the recommended ViT-S+ model takes 17.2 ms per image with TensorRT FP32, including preprocessing, transfers and triplet decoding. Object regions are provided as input.
With YOLO26m detection included, the full pipeline takes 36.5 ms per image. Detection uses PyTorch FP32; ViT-S+ relations use TensorRT FP32.
Warm median latency, batch 1, 35 relation types, 6 sample images; loading and setup excluded.
All models and runtimes · YOLO26, YOLO-World and YOLOE comparison · Paper pipeline benchmarks
The paper evaluates relation prediction across datasets, open vocabularies and spatial reasoning tasks. Selected transfer results for ViT-S+, using ground-truth boxes and one relation per ordered pair:
| Test set | OvSGTR F1@50 | RelateAnything F1@50 |
|---|---|---|
| VG150 | 16.5 | 36.9 |
| PSG | 13.5 | 34.7 |
| IndoorVG | 20.2 | 37.8 |
Both models use the same evaluator and vocabulary. OvSGTR receives object labels; RelateAnything does not. The full results cover stronger baselines including ROBIN-3B, zero-shot controls and the complete evaluation.
Evaluation protocol · Scoring guide
Figure 2 from the paper. Click to enlarge.
A DINOv3 backbone reads the image once. Regions become visual tokens; a pair sampler selects candidate relations; a transformer refines them using the scene. A vocabulary head scores each pair against embeddings of your predicate strings, mixing spatial and semantic features. Swapping that embedding bank changes the vocabulary without retraining the vision model.
- RA-4M: training annotations for open-vocabulary relations. Images are referenced by ID and are not redistributed. Dataset details and generation.
- OV-SGG-Bench: evaluation packs, explicit negatives and calibration files for six complementary axes. Protocol.
- Train: training guide, released configs
and
train.sh. - Evaluate: evaluation guide and
benchmark/. - Generate annotations:
datagen/. Explore the paper's probes:research/. - Contribute: contributor guide. Bug reports, deployment feedback and examples of new uses are welcome in GitHub issues. Include the model, runtime, hardware and a small reproducer when reporting a problem.
The documentation index covers the API, architecture, training, data and deployment. Read evaluation pitfalls before comparing results or changing the scoring path.
- Code: AGPL-3.0-only, with the additional attribution terms in NOTICE. If you run a modified version as a network service, you must offer its users the complete source of that version.
- Commercial license: to use RelateAnything in a product or service without the AGPL obligations, contact the author (Maëlic Neau) through GitHub issues or the email in the paper.
- Model weights: derivatives of Meta DINOv3; see the model cards and DINOv3 license.
- RA-4M annotations: carry the Gemma Terms of Use notice. Source images are not redistributed.
- Demo detectors: ultralytics-derived components use AGPL-3.0 and are excluded from the relation model's release artifacts.
Full third-party notices and credits
@article{neau2026relateanything,
title = {RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs},
author = {Neau, Ma\"elic},
journal = {arXiv preprint arXiv:2609.12552},
eprint = {2609.12552},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2609.12552},
year = {2026}
}