Skip to content
Merged
Show file tree
Hide file tree
Changes from 1 commit
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -151,7 +151,7 @@ uv run python examples/quickstart.py path/to/episode.mcap
For the recommended first run on real egocentric footage, follow
[Evaluate your first egocentric episode](./docs/tutorials/evaluate-an-egocentric-episode.md).
It runs HFlow's local deterministic quality baseline and two hosted semantic
checks together, without requiring an account, API key, or model server.
checks together. Contact us to get access to the hosted API.

Use `uv run hflow --help` to see the CLI. When you are ready to schedule the
same pipeline, continue with the [runtime guide](https://github.com/Hebbian-Robotics/hflow/blob/main/docs/RUNTIME.md). Developers
Expand Down
8 changes: 4 additions & 4 deletions docs/FAQ.md
Original file line number Diff line number Diff line change
Expand Up @@ -68,10 +68,10 @@ where in the recording each problem occurred.

By default nothing leaves your infrastructure: the SDK, the catalog, and every
built-in check run where you run them. The hosted checks are the exception and
are opt-in. When you register one with `HFlowHostedExecution`, the SDK sends
the sampled JPEG frames that check needs to `https://api.hflow.dev` and records
the answers in your catalog alongside every other check. No API key is needed.
Each hosted check version is frozen, so results stay comparable over time, and
are opt-in. Contact us to get access to the hosted API. When you register a
check with `HFlowHostedExecution(base_url=...)`, the SDK sends the sampled JPEG
frames that check needs to that URL and records the answers in your catalog
alongside every other check. Each hosted check version is frozen, so results stay comparable over time, and
the service admits one request at a time per client.

## What does a run produce?
Expand Down
30 changes: 17 additions & 13 deletions docs/how-to/run-build-ai-evaluation.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,8 +24,9 @@ deterministic quality checks, start with

## Run the methodology on an HFlow episode

The example uses HFlow's fixed hosted implementation for both checks by
default. It needs no model configuration, account, or API key.
The example sends both checks to an OpenAI-compatible vision endpoint by
default. HFlow's fixed hosted implementation is also available; contact us to
get access to the hosted API.

To try the checks without supplying your own recording, download a small
egocentric MCAP from Lightwheel's
Expand All @@ -46,9 +47,12 @@ cp "$downloaded_sample_mcap_path" data/build-ai-evaluation/sample.mcap

If Hugging Face requests access, accept the dataset's terms and run
`hf auth login` before downloading. You can then run the pipeline against the
sample:
sample with an unauthenticated local OpenAI-compatible vision server:

```bash
export OPENAI_BASE_URL="http://localhost:8000/v1"
export OPENAI_MODEL="Qwen/Qwen3-VL-8B-Instruct"

uv run --project examples/build_ai_evaluation \
python -m examples.build_ai_evaluation.pipeline \
data/build-ai-evaluation/sample.mcap
Expand All @@ -57,13 +61,12 @@ uv run --project examples/build_ai_evaluation \
Replace that path with any MCAP containing meaningful egocentric footage to
evaluate your own recording.

To use an unauthenticated local OpenAI-compatible vision server instead, select
that execution route and identify its model:
To use HFlow's hosted checks instead, select that execution route and pass the
hosted API URL you received:

```bash
export BUILD_AI_EXECUTION="openai-compatible"
export OPENAI_BASE_URL="http://localhost:8000/v1"
export OPENAI_MODEL="Qwen/Qwen3-VL-8B-Instruct"
export BUILD_AI_EXECUTION="hflow-hosted"
export BUILD_AI_HOSTED_BASE_URL="$HFLOW_HOSTED_BASE_URL"

uv run --project examples/build_ai_evaluation \
python -m examples.build_ai_evaluation.pipeline \
Expand Down Expand Up @@ -107,8 +110,8 @@ different configurations do not claim comparable step versions.
Each check can select its execution independently with
`BUILD_AI_HAND_VISIBILITY_EXECUTION` and
`BUILD_AI_ACTIVE_MANIPULATION_EXECUTION`; each defaults to `BUILD_AI_EXECUTION`
and then to `hflow-hosted`. For example, one check can remain hosted while the
other uses your model:
and then to `openai-compatible`. For example, one check can use the hosted API
while the other uses your model:

```bash
export BUILD_AI_HAND_VISIBILITY_EXECUTION="hflow-hosted"
Expand All @@ -121,7 +124,8 @@ OpenAI-compatible checks can use different services through the
`BUILD_AI_ACTIVE_MANIPULATION_*` overrides.

`HFlowHostedExecution` owns the hosted base URL, check version, and request policy;
its server owns every model setting. `request_timeout_seconds` defaults to 60,
its server owns every model setting. `base_url` has no default: contact us to
get access to the hosted API. `request_timeout_seconds` defaults to 60,
`total_timeout_seconds` to 360, and `max_retries` to five additional attempts.
Tenacity retries transport failures and HTTP 429/502/503/504, respecting numeric
`Retry-After` delays (capped at 120 seconds) or exponential backoff. Authorization
Expand Down Expand Up @@ -165,7 +169,7 @@ per check, so the two checks may use different executions:
```python
hflow.build_ai_vlm_checks.register_hand_visibility(
app,
execution=hflow.build_ai_vlm_checks.HFlowHostedExecution(),
execution=hflow.build_ai_vlm_checks.HFlowHostedExecution(base_url=hosted_api_url),
)
hflow.build_ai_vlm_checks.register_active_manipulation(
app,
Expand All @@ -191,7 +195,7 @@ the same prompt to every frame at that rate instead:
```python
hflow.build_ai_vlm_checks.register_hand_visibility(
app,
execution=hflow.build_ai_vlm_checks.HFlowHostedExecution(),
execution=hflow.build_ai_vlm_checks.HFlowHostedExecution(base_url=hosted_api_url),
sampling=hflow.build_ai_vlm_checks.FrameSampling(fps=1.0, start_s=0.0, end_s=None),
)
```
Expand Down
10 changes: 6 additions & 4 deletions docs/tutorials/evaluate-an-egocentric-episode.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,8 +5,8 @@ point for egocentric video: the default deterministic quality checks plus two
hosted model checks. It demonstrates the useful split in one command:

- HFlow measures recording integrity locally and deterministically.
- HFlow's hosted checks add semantic evidence without requiring an account,
API key, or model server.
- HFlow's hosted checks add semantic evidence without a model server of your
own. Contact us to get access to the hosted API.
- Both kinds of evidence are versioned and recorded through the same pipeline.

The hosted checks are registered explicitly rather than included in
Expand Down Expand Up @@ -50,11 +50,13 @@ The sample has two camera streams, so select its left head camera explicitly:
```bash
uv run python examples/evaluate_episode.py \
data/episode-evaluation/sample.mcap \
--hosted-base-url "$HFLOW_HOSTED_BASE_URL" \
--camera /sensor/camera/head_left/video
```

No model endpoint or API key is required. HFlow processes the episode locally;
the two model checks send only the selected JPEG frame to `api.hflow.dev`.
Set `HFLOW_HOSTED_BASE_URL` to the hosted API URL you received. HFlow
processes the episode locally; the two model checks send only the selected
JPEG frame to that URL.

The first sample frame visibly contains two wearer hands manipulating material.
The hosted results should therefore include:
Expand Down
23 changes: 14 additions & 9 deletions examples/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -172,8 +172,9 @@ Code: [`measurement_distribution.py`](./measurement_distribution.py)
**Use it for:** seeing HFlow's default deterministic checks and hosted semantic
checks work together on real egocentric footage.

**Prerequisites:** the normal development environment, the `hf` CLI, and
network access. The hosted checks require no account, API key, or model server.
**Prerequisites:** the normal development environment, the `hf` CLI, network
access, and a hosted API URL in `HFLOW_HOSTED_BASE_URL`. Contact us to get
access to the hosted API.

```bash
mkdir -p data/episode-evaluation
Expand All @@ -187,6 +188,7 @@ cp "$downloaded_sample_mcap_path" data/episode-evaluation/sample.mcap

uv run python examples/evaluate_episode.py \
data/episode-evaluation/sample.mcap \
--hosted-base-url "$HFLOW_HOSTED_BASE_URL" \
--camera /sensor/camera/head_left/video
```

Expand Down Expand Up @@ -229,9 +231,9 @@ Code: [`openai_vision/pipeline.py`](./openai_vision/pipeline.py)
active-manipulation methodology as HFlow checks, then replaying their
Egocentric-10K or Egocentric-100K inputs to compare models and prompts.

**Prerequisites:** network access and the `hf` CLI for the sample recording.
The default pipeline route uses HFlow's hosted checks without an account or API
key. The pinned evaluation replay additionally needs sufficient disk for the
**Prerequisites:** network access, the `hf` CLI for the sample recording, and an
OpenAI-compatible vision endpoint (or access to HFlow's hosted API; contact us).
The pinned evaluation replay additionally needs sufficient disk for the
selected Parquet files. Each frame makes two external check calls; a configured
model provider may charge for them.

Expand All @@ -245,6 +247,9 @@ downloaded_sample_mcap_path="$(
)"
cp "$downloaded_sample_mcap_path" data/build-ai-evaluation/sample.mcap

export OPENAI_BASE_URL="http://localhost:8000/v1"
export OPENAI_MODEL="Qwen/Qwen3-VL-8B-Instruct"

uv run --project examples/build_ai_evaluation \
python -m examples.build_ai_evaluation.pipeline \
data/build-ai-evaluation/sample.mcap
Expand All @@ -256,12 +261,12 @@ identity, and any available model usage as `hflow.CheckResult`
measurements. An episode path is required because generated test-pattern
footage cannot meaningfully demonstrate either judgment.

To use a local OpenAI-compatible model server instead of HFlow's hosted checks:
To use HFlow's hosted checks instead (contact us to get access to the hosted
API):

```bash
export BUILD_AI_EXECUTION="openai-compatible"
export OPENAI_BASE_URL="http://localhost:8000/v1"
export OPENAI_MODEL="Qwen/Qwen3-VL-8B-Instruct"
export BUILD_AI_EXECUTION="hflow-hosted"
export BUILD_AI_HOSTED_BASE_URL="$HFLOW_HOSTED_BASE_URL"

uv run --project examples/build_ai_evaluation \
python -m examples.build_ai_evaluation.pipeline \
Expand Down
18 changes: 11 additions & 7 deletions examples/build_ai_evaluation/pipeline.py
Original file line number Diff line number Diff line change
@@ -1,8 +1,9 @@
"""Run Build AI's published single-frame checks on HFlow episodes.

The example uses HFlow's fixed hosted checks by default, without an API key.
Set ``BUILD_AI_EXECUTION=openai-compatible`` to use a local or third-party
OpenAI-compatible vision endpoint instead.
The example sends each frame to an OpenAI-compatible vision endpoint, local or
third-party. Set ``BUILD_AI_EXECUTION=hflow-hosted`` and
``BUILD_AI_HOSTED_BASE_URL`` to use HFlow's hosted checks instead; contact us to
get access to the hosted API.
"""

from __future__ import annotations
Expand All @@ -18,8 +19,9 @@
DEFAULT_HFLOW_DATA_ROOT = Path("data/build-ai-evaluation/hflow")
DEFAULT_OPENAI_COMPATIBLE_BASE_URL = "http://localhost:8000/v1"
DEFAULT_MODEL_NAME = "model-not-configured"
DEFAULT_EXECUTION_NAME = "hflow-hosted"
HFLOW_HOSTED_EXECUTION_NAME = "hflow-hosted"
OPENAI_COMPATIBLE_EXECUTION_NAME = "openai-compatible"
DEFAULT_EXECUTION_NAME = OPENAI_COMPATIBLE_EXECUTION_NAME


def _optional_float_environment_variable(name: str, environment: Mapping[str, str]) -> float | None:
Expand All @@ -41,8 +43,10 @@ def _execution_from_environment(
check_execution_environment_variable,
environment.get("BUILD_AI_EXECUTION", DEFAULT_EXECUTION_NAME),
)
if execution_name == DEFAULT_EXECUTION_NAME:
return hflow.build_ai_vlm_checks.HFlowHostedExecution()
if execution_name == HFLOW_HOSTED_EXECUTION_NAME:
return hflow.build_ai_vlm_checks.HFlowHostedExecution(
base_url=environment.get("BUILD_AI_HOSTED_BASE_URL")
)
if execution_name == OPENAI_COMPATIBLE_EXECUTION_NAME:
return hflow.build_ai_vlm_checks.OpenAICompatibleExecution(
endpoint=environment.get(
Expand All @@ -69,7 +73,7 @@ def _execution_from_environment(
)
raise ValueError(
f"{check_execution_environment_variable} (or BUILD_AI_EXECUTION) must be "
f"{DEFAULT_EXECUTION_NAME!r} or {OPENAI_COMPATIBLE_EXECUTION_NAME!r}, "
f"{HFLOW_HOSTED_EXECUTION_NAME!r} or {OPENAI_COMPATIBLE_EXECUTION_NAME!r}, "
f"got {execution_name!r}"
)

Expand Down
23 changes: 18 additions & 5 deletions examples/build_ai_evaluation/tests/test_evaluation.py
Original file line number Diff line number Diff line change
Expand Up @@ -267,23 +267,36 @@ def test_build_ai_pipeline_registers_both_judgments_as_hflow_checks() -> None:
assert all(check.requires == frozenset({"vision-model"}) for check in checks_by_name.values())


def test_build_ai_pipeline_defaults_to_hosted_and_can_select_openai_compatible_execution() -> None:
hosted_execution = _execution_from_environment("BUILD_AI_HAND_VISIBILITY", {})
def test_build_ai_pipeline_defaults_to_openai_compatible_and_can_select_hosted_execution() -> None:
openai_compatible_execution = _execution_from_environment(
"BUILD_AI_HAND_VISIBILITY",
{
"BUILD_AI_EXECUTION": "hflow-hosted",
"BUILD_AI_HAND_VISIBILITY_EXECUTION": "openai-compatible",
"BUILD_AI_HAND_VISIBILITY_BASE_URL": "http://localhost:8000/v1",
"BUILD_AI_HAND_VISIBILITY_MODEL": "local-vision-model",
},
)
hosted_execution = _execution_from_environment(
"BUILD_AI_HAND_VISIBILITY",
{
"BUILD_AI_HAND_VISIBILITY_EXECUTION": "hflow-hosted",
"BUILD_AI_HOSTED_BASE_URL": "https://hosted.example",
},
)

assert hosted_execution == hflow.build_ai_vlm_checks.HFlowHostedExecution()
assert openai_compatible_execution == hflow.build_ai_vlm_checks.OpenAICompatibleExecution(
endpoint="http://localhost:8000/v1",
model="local-vision-model",
)
assert hosted_execution == hflow.build_ai_vlm_checks.HFlowHostedExecution(
base_url="https://hosted.example"
)


def test_build_ai_pipeline_hosted_execution_without_a_url_asks_the_user_to_contact_us() -> None:
with pytest.raises(ValueError, match="contact us"):
_execution_from_environment(
"BUILD_AI_HAND_VISIBILITY", {"BUILD_AI_EXECUTION": "hflow-hosted"}
)


def test_build_ai_pipeline_requires_an_episode_with_meaningful_footage() -> None:
Expand Down
13 changes: 11 additions & 2 deletions examples/evaluate_episode.py
Original file line number Diff line number Diff line change
Expand Up @@ -2,10 +2,12 @@

The normal ``hflow.App`` baseline stays enabled. Two explicitly registered
hosted checks add Build AI's hand-visibility and active-manipulation judgments.
Contact us to get access to the hosted API.

Run from the repository root::

uv run python examples/evaluate_episode.py episode.mcap [--camera TOPIC]
uv run python examples/evaluate_episode.py episode.mcap \\
--hosted-base-url URL [--camera TOPIC]
"""

from __future__ import annotations
Expand All @@ -22,11 +24,12 @@
def build_application(
*,
data_root: Path,
hosted_base_url: str,
camera: str | None,
frame_time_seconds: float,
) -> hflow.App:
application = hflow.App("episode-evaluation", data_root=data_root)
hosted_execution = hflow.build_ai_vlm_checks.HFlowHostedExecution()
hosted_execution = hflow.build_ai_vlm_checks.HFlowHostedExecution(base_url=hosted_base_url)
hflow.build_ai_vlm_checks.register_hand_visibility(
application,
execution=hosted_execution,
Expand All @@ -45,6 +48,11 @@ def build_application(
def argument_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("episode", type=Path, help="MCAP episode with egocentric video")
parser.add_argument(
"--hosted-base-url",
required=True,
help="HFlow hosted API URL; contact us to get access",
)
Comment thread
greptile-apps[bot] marked this conversation as resolved.
parser.add_argument(
"--camera",
help="camera topic to evaluate; required when the episode has multiple cameras",
Expand All @@ -68,6 +76,7 @@ def main() -> None:
arguments = argument_parser().parse_args()
application = build_application(
data_root=arguments.data_root,
hosted_base_url=arguments.hosted_base_url,
camera=arguments.camera,
frame_time_seconds=arguments.frame_time_seconds,
)
Expand Down
10 changes: 8 additions & 2 deletions src/hflow/build_ai_vlm_checks.py
Original file line number Diff line number Diff line change
Expand Up @@ -125,7 +125,6 @@

BUILD_AI_HAND_VISIBILITY_CHECK_NAME = "build_ai_hand_visibility"
BUILD_AI_ACTIVE_MANIPULATION_CHECK_NAME = "build_ai_active_manipulation"
DEFAULT_HFLOW_HOSTED_BASE_URL = "https://api.hflow.dev"

_HFLOW_HOSTED_TRANSPORT_VERSION = 1
_DEFAULT_HFLOW_HOSTED_CHECK_VERSION = 1
Expand Down Expand Up @@ -301,7 +300,8 @@ class HFlowHostedExecution:
Build AI's prompts or results.
"""

base_url: str = DEFAULT_HFLOW_HOSTED_BASE_URL
# No default: access to the hosted service is granted on request.
base_url: str | None = None
check_version: int = _DEFAULT_HFLOW_HOSTED_CHECK_VERSION
request_timeout_seconds: float = 60.0
total_timeout_seconds: float = 360.0
Expand All @@ -312,6 +312,10 @@ class HFlowHostedExecution:
max_retries: int = 5

def __post_init__(self) -> None:
if self.base_url is None:
raise ValueError(
"HFlowHostedExecution needs base_url: contact us to get access to the hosted API"
)
_require_absolute_http_url(self.base_url, name="base_url")
require_non_negative_int(self.max_retries, "max_retries")
parsed_base_url = urlsplit(self.base_url)
Expand Down Expand Up @@ -806,6 +810,8 @@ def _check_name_for_task(task: EvaluationTask) -> str:

def _hosted_check_endpoint(execution: HFlowHostedExecution, task: EvaluationTask) -> str:
check_name = _check_name_for_task(task)
# __post_init__ refuses a missing base_url.
assert execution.base_url is not None
return (
f"{execution.base_url.rstrip('/')}/v{_HFLOW_HOSTED_TRANSPORT_VERSION}/checks/"
f"{check_name}/versions/{execution.check_version}/evaluate"
Expand Down
Loading
Loading