Evaluation Module#
Note: This module is still under active development and may see changes or adjustments in the future.
This directory contains the LoongForge-Embodied offline evaluation module. It runs benchmark clients and model policy servers as separate processes connected by a WebSocket/msgpack-numpy RPC protocol.
Architecture#
The benchmark client and the policy server run as separate processes:
Benchmark env (client) Policy server
├─ Adapter ├─ ModelFactory
│ env obs → canonical observation │ loads the model, wraps predict_action
├─ PayloadBuilder ├─ GenericPredictActionPolicy
│ canonical → predict_action kwargs │ RPC / chunk cache / shape check
└─ ActionDecoder └─ model.predict_action
raw chunk → env action model inference + normalization
◄──────── WebSocket + msgpack-numpy RPC ────────►
PolicyRequest / Response
Every benchmark drives the same four-stage chain: Adapter.obs_to_canonical → PayloadBuilder.build → model.predict_action (over RPC) → ActionDecoder. Details for code contributors: the evaluation protocol and the model integration guide.
Docs#
Doc |
Audience |
Content |
|---|---|---|
eval users |
Supported models/benchmarks, quick start, config reference, outputs, troubleshooting |
|
eval users |
Per-benchmark reproduction guides (env setup, run, verification) |
|
code contributors |
New-model integration checklist + |
|
code contributors |
Architecture, components, data protocol, PolicyClient interface |
|
eval users |
Verified per-benchmark env version records (install per the official benchmark homepages) |
Quick Start#
The example below runs LIBERO with pi05:
Set up the LIBERO environment — please refer to the official LIBERO repository for installation; for the eval client deps and common issues, check the LIBERO guide Environment setup section; the verified environment version lists are in benchmark_envs.md.
Get the weights and edit the config — download lerobot/pi05_libero_finetuned_v044, then fill the
/path/to/...placeholders inexamples/embodied/pi05/eval/configs/libero/object_smoke.yaml. A field-by-field example: the user guide §2 Quick start; the config layout: user guide §3 Configuration reference.Run — execute the script inside the benchmark environment:
cd /path/to/LoongForge examples/embodied/pi05/eval/run_libero_eval.sh
The LIBERO simulator runs in the benchmark environment; the policy server is launched from the YAML server.python field, pointing at the LoongForge environment. Other models/benchmarks: the user guide and the benchmark pages.
Supported Matrix#
LIBERO |
CALVIN |
SimplerEnv (WidowX) |
RoboTwin |
ManiSkill |
|
|---|---|---|---|---|---|
pi05 |
✅ task success (lerobot/pi05_libero_finetuned_v044) |
connectivity only |
connectivity only |
✅ task success (motus-robotics/pi0.5_robotwin2) |
connectivity only |
xvla |
✅ task success (2toINF/X-VLA-LIBERO) |
connectivity only |
✅ task success (2toINF/X-VLA-WidowX) |
✅ task success (2toINF/X-VLA-RoboTwin2) |
connectivity only |
GR00T-N1.6 |
✅ task success (0xAnkitSingh/GR00T-N1.6-LIBERO) |
— |
✅ task success (nvidia/GR00T-N1.6-bridge) |
— |
— |
Task success: at least one episode passed the benchmark’s official success criterion.
Weights: parentheses show the Hugging Face weight (
org/name) that achieved the run.Connectivity only: the pipeline runs with
random_init: trueand no score — either no domain weights are released, or the benchmark assets block a full run (e.g. xvla CALVIN: the weights are public, but the official online rollout needs the original-format validation dataset).—: not supported yet — coming soon.
Detailed results: see the per-benchmark benchmark pages.
Model Interface#
LoongForge model servers share a predict_action(images, instructions, state=None, dataset_stats=None) interface instead of a bespoke policy adapter per model. Client-side payload assembly lives in the per-model PayloadBuilder; env-side action decoding lives in the ActionDecoder; normalization stays inside the model’s predict_action(). Architecture and protocol: the evaluation protocol. Step-by-step integration checklist: the model integration guide.